Category: SCOSYA
Week Beginning 3rd August 2026
I worked a total of four days over the past two weeks, and was on holiday for the remainder. During this time I had a meeting to further discuss a place-names related AHRC proposal I’m involved with. I can’t really say much more about it at this stage, but the proposal is coming together. I also had to spend some time working with IT Support and Luca to figure out why our local server kept going offline repeatedly. It looks like this was caused by the server getting swamped by requests from one particular source (almost certainly bot or AI) and thankfully IT Support were able to block this, after which the server was stable again. It’s something we’re going to have to keep looking out for in future. I also spent a bit of time working with Luca to get automatic WordPress updates working on our local server, as the way sites had been set up meant that the setting was not working. Luca managed to find a solution to this, which is really great.
I also spent a bit more time working on the Burns Supper Map, creating a record for it on this site (see https://digital-humanities.glasgow.ac.uk/project/?id=156), adding more suppers that had been submitted via the survey and making some requested edits to existing suppers. I also set up access to Google Analytics for the other two members of the project team.
In addition, I investigated an issue with the Scots Syntax Atlas after someone suggested that the linguists’ atlas was looking somewhat blurry. I managed to figure out why this might be the case, although I’m not entirely sure whether this is a new issue or if the markers always looked like that. I contacted the project PI and suggested a couple of updates, but I haven’t heard back yet so will need to wait and see what she says.
I spent most of the remainder of my time working on updates to our new test interface for Dictionaries of the Scots Language and working through the list of outstanding items for the DOST Auld Laws project. For the DSL I completed the updates to the bibliography page that I began working on a couple of weeks ago. I implemented pagination of the entries associated with bibliographical items, with navigation bars appearing above and below the entries, with 20 appearing per page and ‘jump to page’ buttons also appearing, just like with the search results. This works pretty well, but is somewhat cumbersome for someone like Sir Walter Scott, who is referenced in 2861 entries, split over 144 pages. We don’t have this issue with the search results are these are capped at 500 (25 pages) so we might need to think of other ways of handling this.
I also ensured that headword searches that don’t yield any results automatically perform a fulltext search for the term supplied. This works on the live site, but only when both dictionaries are selected. With the new site there are separate quick searches for SND and DOST so the additional search wasn’t being triggered. It is now, as is the advanced headword search for both dictionaries.
I also made tweaks to the DSL’s new regional map based on feedback I’d received – adding in some content where we previously had placeholder text and ensuring the ‘About’ popup didn’t disappear off the bottom of smaller screens and a few other small updates. I then began to look at the ancillary pages and how we can make them look a bit nicer. I spent a bit of time on the ‘Word of the Week’ page and liaised with William Ashford, who is responsible for such content about this, and further updates that were going to make to the ancillary content closer to the launch date of the new site (which will hopefully be in November).
For the DOST Auld Laws project I added the top navigation bar that will link the site in with the SCOTS Corpus and CMSW. I added in copyright information and added facilities to download page images and the XML files for each document. These being up a pop-up asking for users to abide by the license before leading to the actual content, which hopefully won’t be too annoying. I also added in the ‘cite’ popup to all document pages, which took a little time to implement, and added in Google Analytics. I also made the image thumbnails on the document overview pages smaller and placed them in a collapsible section that is closed by default, plus I removed the introduction to the documents page, as this will be covered by the homepage.
I also removed some pages that didn’t have content (e.g. blank pages) from the beginning and end of some of the documents and I added a feature to turn off and on the line highlighting feature. The highlighting feature allows the user to click on a line of text in the image or text for a page and for that line to be highlighted in both the image and the text, which is pretty nice. Unfortunately the line highlighting gets in the way of the image viewer’s zoom and pan functionality on touchscreens, making it a somewhat unreliable and frustrating experience. This new feature removes the option to ‘click’ on a line, meaning pointer events are not intercepted and make their way reliably through to the image viewer, which works much more smoothly.
Also this week I had an email conversation about user feedback and walkthough videos for the STAR resources and booked my accommodation for the DHC conference in Sheffield. Next week I’m back in Glasgow and back working a full week, with summer holidays all over.
Week Beginning 19th January 2026
I worked on many different projects this week, but the one I spent the most time on was the Dictionaries of the Scots Language. I’ve not done much work for the DSL since the intensive period I spent developing the new website interface and deploying it on our test server ahead of the face-to-face meeting in mid-November. I had a list of further updates I needed to make following on from this meeting, but I needed to work on other projects since then and hadn’t got around to it. I’d also received a number of emails about changes to the presentation of entries reflecting the structural changes to the entry XML that I’d put to one side.
On Wednesday I had an online call scheduled with the DSL team to discuss the new front-end and it seemed like a good opportunity to get back to grips with all of my outstanding DSL tasks. This mainly involved making updates to the XSLT on our test server to tweak the layout of various items in the new entry XML structure, such as adding commas between tags when they are rendered, ensuring certain tags or attributes that weren’t getting rendered before appeared in the generated HTML, updating the styles of certain elements like the content warning labels, fixing a few bugs such as the ‘sticky’ heading not displaying in certain circumstances. There were at least 20 such items that needed investigating, fixing and testing, so this took quite some time to work through, but I managed to complete it all during the course of the week.
The meeting itself was very useful and as always it was good to catch up with some of the other DSL team members. There’s going to be a big push towards getting the new website interface ready for publication this year, and I’m obviously going to be involved in this process. I already have a number of items I need to sort out with the new interface and I’ll try and get started on these over the coming weeks.
I also spent a bit of time this week working for the Anglo-Norman Dictionary, investigating a strange occurrence with the publication of updates to entries, which turned out to be a user rather than a system issue, reinstating the links out from entries to the DMF dictionary, as their website is now properly back online again, and tweaking the wording of the quick search and ‘jump to entry’ text throughout the site.
I also did small amounts of work for several other projects, such as updating the licensing statements across the Seeing Speech and Star sites, fixing an issue with the ‘download song’ facilities on the Editing Robert Burns site, sorting an issue with the HiMuJe Malabar site, submitting my expenses from the Zurich workshop, exporting some SCOSYA data for Jennifer Smith, helping to sort out an issue with the Helsinki Corpus, and having a conversation with Clara Cohen about a new proposal she’s putting together. I also made some further updates to the VARICS look-up system, adding in some introductory text, some further references, and reworking the measurement processing so that when a red or amber result is given new textual sections about what this means and what the next steps should be appear underneath in collapsible accordion sections.
Also this week I had a meeting with the Burns Supper Map team to discuss the data that is now coming in and how and when I should start working on a new interactive map to visualise it. I also put in a request for a new subdomain for the project that was set up by the end of the week. Next week I’ll probably write a brief specification document for the front-end.
My final project of the week was the Place-names of Armagh project, for which I started working with some existing place-name data for the area. There are around 230 place-names and several thousand historical forms and I spent quite some time researching how the data was structured and how it might be mapped onto the Glasgow place-names system. This included analysing the geospatial data, including shapefiles for Townlands and what I though was Parishes (but actually turned out to be the same as for Townlands). I had hoped to be able to import the data by the end of the week, but my analysis of the data raised a lot of questions that still need to be addressed, and I’ll need to continue with this next week.
Week Beginning 3rd November 2025
I spent pretty much the entirety of this week continuing to develop the new interface for the Dictionaries of the Scots Language website, applying the Bootstrap-based mock-up I’d created many months ago to an instance of the actual DSL website running on my laptop. I can’t really go into too much detail about the new interface or provide any screenshots at this stage, but it’s been a pretty intensive process as every aspect of the old interface needs to be changed and various parts of it need to be integrated with WordPress, for example making widgets and ensuring the new layout works with different templates.
I managed to complete the bulk of the work this week (although this did include working several hours over the weekend too), in preparation for next week’s face-to-face DSL team meeting. This included the search results pages, the advanced search page, the dictionary entry page and the bibliography page. This may not seem like a very long list, but there was a huge amount of work to do on each of these pages, such as implementing the site panel for the entry page that features the dictionary browser, the search results browser and a new ‘entry log’ that keeps a record of entries the user has looked at during their session. I reckon the new interface is looks really good, and is a massive improvement on the live site, although there will inevitably be many further changes to be made before anything goes live.
I still need to complete the new top-level ‘About’ page, which acts as a large menu page, plus ensure that all regular WordPress pages work with the new interface and include the quick search. I’m hoping to finish these things off and then apply the interface to our online test instance of the site ahead of Wednesday’s meeting next week.
Also this week I spent a bit more time preparing for a talk about Speak For Yersel and Jennifer Smith and I were scheduled to give at the University of Edinburgh the week after next. However, later in the week we heard from the organisers that the University will be on strike when our talk is scheduled and we therefore reached a decision to cancel. It’s possible that we’ll be able to reschedule, as we are not directly involved in the strike action, but we’ll just need to see.
Also this week I created an initial version of a website for Henry Ivry’s project and contacted researcher Jenny Buckley with some further information about the processing of historical newspapers that might be of use for her project. I also made a small update to the Speech Star resource and had an email conversation with Eleanor Lawson about access restrictions for the resources data.
I participated in an online meeting regarding sharing the SCOSYA data with the Mozilla Foundation this week, and I also had a meeting with Pauline Mackay and Cleo O’Callaghan Yeoman to discuss a new phase for the Interactive Map of Burns Suppers. I subsequently spent a bit of time reviewing some materials for the site. Finally, I exported some data from the Historical Thesaurus that we’re going to share with another project.
Week Beginning 11th August 2025
I was on holiday for most of last week and a little of this week, so I’m joining the weeks together in this update. During this time I worked on several different projects and communicated with a number of people about new research projects.
Whilst at the DH2025 conference in Lisbon Jennifer Smith alerted me to a problem with the sending of emails on the Scots Syntax Atlas site. We have a couple of contact forms (including one for requesting the audio files generated by the project) and neither were working. I’d also spotted that the two-factor authentication plugin used when logging into the WordPress admin interface was failing to send the emails containing the log-in codes. This took an awfully long time to sort out as it was unclear what was causing the problem or whose responsibility it was to fix it, compounded by the fact that the site is hosted by an external company. I’d noticed that several other of our externally hosted sites had the same issue with sending the 2FA emails so it appeared to be more widespread than just one site.
I spoke to Andrew McHugh, the head of RCaaS in IT Services and he suggested contacting the third party hosting company’s support. I had a long and useful chat with their support people, who suggested that we’d need to update the DNS setting relating to mail for the scotssyntaxatlas.ac.uk domain (which Glasgow controls) to point to the hosting company’s email systems. This would then enable me to set up SMTP for email for the domain and hopefully get the emails sending again. I therefore submitted a helpdesk request to Glasgow’s IT people asking them to update the record.
At the same time, I wasn’t convinced that this would solve the problem, and it was very strange that the 2FA issue was only affecting certain sites and not others. I realised that it was only the sites that had their own domain rather than being a subdomain of Glasgow that were being affected. I spoke to Luca who confirmed that this was also the case for one of his sites that had its own domain. It was also strange that the WordFence summary emails for these sites were getting sent, which suggested that emails were somehow working. I eventually realised that the 2FA emails were being sent from the email address ‘wp2fa@domain’ while the WordFence emails were being sent from ‘wordpress@domain’. There was an optin in the 2FA plugin to change the email address its messages were sent from so I set this to ‘wordpress@domain’ and… the messages came through.
It turns out that the University must have blacklisted the ‘wp2fa@domain’ emails, meaning all messages from these addresses were being blocked. So there was never an issue with the sending of emails for the domains – the messages were being sent but were being blocked and changing the email address to ‘wordpress@domain’ allowed them to get through.
This meant that what I thought was one problem was actually two, as the contact forms on the Scots Syntax Atlas site were still broken. I’d noticed previously that there appeared to be an issue with the contact forms’ use of Google ReCaptcha, but when I’d disabled this the forms still failed to send. Now I knew the problem was definitely related to the contact form plugin I could investigate this further. I made a copy of the contact form details and then completely deleted the contact form plugin, then reinstalled it and set up the forms again without any involvement with ReCaptcha. Thankfully this fixed the issue and the contact forms successfully sent their emails. It does mean that there may be a lot more spam sent to the email address that receives the messages, and we’ll just need to keep an eye on this.
Also during this period I continued to work with the Fife Place-names data. Previously I’d written a script to add latitude and longitude to all of the place-name data (more than 3000 records) that previously only had grid references. I’d run into some issues with some of the grid references, as discussed previously, and this meant researching the correct grid references to use, which took a bit of time. I managed to complete this process during these two weeks, successfully adding latitude and longitude to all but 50 of the records. These 50 records are mostly for obsolete place-names that don’t have a specific location, and I’ve emailed the project PI Simon Saylor to ask about these.
During this time I also set up a new project website for Gavin Miller and migrated the emblems websites over to a new server in an attempt to ensure the sites don’t use up all of the server’s resources, which was occasionally happening with the previous hosting arrangement. I also made some further updates to the Robert Fegusson online exhibition, based on feedback and further data from Amy Wilcockson.
Most of the remainder of my time was spent working for the Dictionaries of the Scots Language. I had an email conversation with Ann regarding the new part of speech tagging system and a further email conversation with Becca, William and Vasilis about the new structure of the DSL entry XML files. I now have some sample data and some helpful explanatory notes about the new structure so an upcoming task will be to ensure the DSL’s online systems can work with this new structure.
I also had an online meeting with Vasilis, Becca and William on Friday to discuss the new region and dialect area maps that will be added to the DSL website. This was a very useful meeting and we went through Vasilis’ prototype and discussed how this could be turned into a resource that would work on mobile devices as well as larger monitors and how we could make access to the large amount of overlapping spatial data a little less confusing. I then began writing a specification document for the new feature and will continue with this next week.
Also this week I had email conversations with Emanuele Scieri of Theology and Mícheál Ó Mainnín, a place-names scholar at Queen’s University Belfast about new projects that I may be involved with. I can’t say much more about these for now, but I’ll be having online meetings to discuss their proposals in the next week.
Week Beginning 12th May 2025
This week finally saw the launch of one of the new ‘map first’ interfaces for the various place-names projects I’ve been working on. I initially created the interface for the Iona place-name project and then applied it to Ayrshire, Kirkcudbrightshire, Berwickshire and Nairnshire, but none of the projects were quite ready to publicly launch the new interface. After having a char with Carole Hough last week I got the go-ahead to go live with the new Berwickshire map, and it can now be found here: https://berwickshire-placenames.glasgow.ac.uk/map with the textual list of Berwickshire place-names here: https://berwickshire-placenames.glasgow.ac.uk/list-of-berwickshire-placenames/.
It’s great to have the new interface publicly available at last, although I did have a number of last-minute tweaks to make as I worked on it. I’ve updated the legend slightly, adding in some explanatory text and ensuring that the text is hidden when the ‘hide legend’ button is pressed – previously the text remained visible, which was a bit pointless. I also spotted that the legend was not being set to hidden be default on narrow screens, even though it should have been, so I fixed this. Another issue I spotted was that the tooltips that display when the legend categorisation is by language would stop working whenever the ‘select all’ option was pressed, and I needed to update the code to ensure the tooltips were reinitialised. With these changes in place I also needed to apply them to the other place-names maps to keep things consistent.
Launching the new map wasn’t the end of the work, however. Berwickshire already had a live map resource, plus search and textual browse facilities. I needed to ensure that the quick search on the main resource website connected through to the new map’s search results and that all links / bookmarks / citations to the old interface redirected through to the appropriate section of the new interface, where possible. So for example, the old element glossary page redirects to the new element glossary popup in the new map interface: https://berwickshire-placenames.glasgow.ac.uk/place-names/?p=element-glossary. I think I’ve caught all of the old links and have ensured the necessary redirects work.
Also this week I had a meeting with Thomas and Sofia regarding the map interface for Iona, which will apparently be publicly launched on the 9th of June. We discussed some of the outstanding tasks they would like to see completed before the launch and I’m going to have a fair amount to do before then. For example, they would like the element glossary to have a ‘jump to letter’ feature, allowing the user to immediately scroll to elements beginning with a selected letter. We also discussed new ‘thematic maps’ based on static maps Sofia had created for a recent event on Iona. These maps showed place-names grouped in more fine-grained ways that are not covered by the classification options we currently have available – for example place-names featuring animal names. I suggested that I could make an interface where such thematic maps could be created in the content management system and displayed in a new menu section in the front end. This would consist of a title and description for the map, and a list of markers that should appear on the map.
I also demonstrated the ‘story map’ interface I’d created for the Scots Syntax Atlas (go here and select ‘Stories behind the examples’: https://scotssyntaxatlas.ac.uk/atlas/). This feature consists of a series slides a user can navigate through, each of which can have a different view of the map, which may feature different data and different zoom level, thus guiding the user through a particular story the map tells. I gave an example of how this could be used for iona to have a story of place-names featuring animals, with different slides showing markers featuring elements from different languages. Thomas and Sofia really liked this idea and would like to implement it, but while it might be possible for me to implement the simpler thematic maps before the 9th of June, it’s unlikely that I’d be able to get the ‘story’ approach in place. It’s something we’ll probably consider after the launch.
Continuing with place-names, I had an email conversation with Alasdair Whyte regarding his Mull / Ulva place-name data. He’s currently working towards a published volume and wanted me to add a facility to the content management system to enable him to categorise place-names by volume. I therefore updated the CMS to add in the volume field, and this now appears in the ‘add’ and ‘edit’ place forms. I also updated the ‘Browse’ page to add in volume to the search options and as a table heading for the listed data and updated the ‘Export place-name data for publication’ page in ‘Tools’ to add in a volume selection option. Thomas also contacted me this week to ask for an update to the ‘Export’ facility for Iona (adding in the ‘translation’ field) so I did this too. I also had a chat with Alasdair about the front-end for his Mull / Ulva data, as this will need to be a bit different from the other resources.
Also this week I finally managed to get my ticket sorted for the Digital Humanities event in Glasgow in June, with the help of Emma McCluskey. I also made some updates to my presentation for the DH event in Lisbon based on some very helpful feedback from Jennifer, and also added in some new publications to the Scots Syntax Atlas resource.
I also spent a bit of time working on the Dictionaries of the Scots Language. I’d been contacted last week by editor Vasilis Karaiskos about the regions data that he has been working on, and the new maps he has been developing to display these regions. I spent some time going through the interface he had developed as a proof of concept, figuring out how it all works and giving some feedback. It will be really great to get more information about the geographical regions into the website and the maps are an excellent starting point. We’re hopefully going to have an online call in the next few weeks to discuss things further.
The rest of my week was devoted to working on the new versioning system for the Books and Borrowing project. I need to update the structure of the Solr index and the other cache files, and also the scripts that generate this data, and this is my next step. As I feared, it’s proving to be an awfully tricky and time-consuming update to implement, and while I did make progress there is still a huger amount to do. I’ll be continuing with this monumental task next week.
Week Beginning 18th September 2023
On Monday and Tuesday this week I participated in the UCU strike action. On my return to work on Wednesday I focussed on writing a Data Management Plan for Jennifer Smith’s ESRC proposal that uses some of the SCOSYA data. After a few follow-up conversations I completed a version of the plan that Jennifer was happy with. I informed her that I’d be happy to help out with any further changes or discussions, but other than that my involvement is now complete.
I spent a fair bit of the remainder of the week trying to fix an old resource. I created the House of Fraser Archive site (https://housefraserarchive.ac.uk/) with Graeme Cannon more than twelve years ago, with Graeme doing the XML parts via an eXist-DB system and me doing the interface and all of the parts that processed and displayed data returned from eXist. Unfortunately the server the site was running on had to be taken offline and the system moved elsewhere. A newer version of eXist was required and the old libraries that were used to connect to the XML database would no longer work. I figured out a way to connect via an alternative method, but this then returned the data in a different structure. This meant I needed to update every page of the site that processed data to not only update the way the system was queried but also update the way the returned data was handled. This took quite a lot of time but I managed to get all of the ‘browse’ options plus the display of records, tags and images working. The only thing I couldn’t get to work was the search, as this seems to use further libraries that are no longer available. I the issue is structuring the query to work with eXist, but I am not much of an expert with eXist and I’m not really sure how to untangle things. I’ve asked Luca if he could have a look at it, as he’s use eXist a lot more than I have. I’ve not heard back from him yet, but hopefully we’ll manage to get the search working, otherwise we may have to remove the search and get people to rely on the browse functions to access the data instead.
For the rest of the week I returned to working on the Books and Borrowing project. One thing on my ‘to do’ list is to sort out the API. There are a few endpoints that I haven’t documented yet, plus the existing documentation and structuring of the API could be improved. I spent some time adding in a license statement and a ‘table of contents’ that lists all endpoints. I’m currently in the middle of adding in the missing endpoint descriptions. After that I’ll need to ensure the examples given all work and make sense and then I need to ensure the CSV output works properly for all data types. I’m fairly certain that some data held in arrays will not output properly as CSV at the moment and this definitely needs sorted.
Week Beginning 11th September 2023
I spent a fair amount of time this week preparing for my PDR session – bringing together information about what I’ve done over the past year and filling out the necessary form. I also had a meeting with Jennifer Smith to discuss an ESRC proposal she’s putting together using some of the data from the SCOSYA project and then spent some further time after the meeting researching some tools the project might use and reading the Case for Support.
I also spent a bit of time working for the Anglo-Norman Dictionary, updating the XSLT file to better handle varlists in citations. So for example instead of:
( MS: s.xiiiex ) Satureia: (A6) gallice savoroye (var. saveray (A9) MS: c.1300 ; saveroy (A12) MS: s.xiii4/4 ; savoreie (B3) MS: s.xiv4/4 ; savoré (C35) MS: s.xv ) Plant Names 230
we’d have:
( MS: s.xiiiex ) Satureia: (A6) gallice savoroye (var. (A9: c.1300) saveray; (A12: s.xiii4/4) saveroy; (B3: s.xiv4/4) savoreie; (C35: s.xv) savoré) Plant Names 230
I completed an initial version of the update using test files and after discussions with the editor Geert and a few minor tweaks the update went live on Wednesday.
I also spent a bit of time working to fix the House of Fraser Archive website, which I created with Graeme Cannon many moons ago. It uses an eXist XML database but needed to be migrated to a new server with more modern versions due to security issues. I spent some time figuring out how to connect to the new eXist database and had just managed to find a solution when the server went down and I was unable to access it. It was still offline at the end of the week, which is a bit frustrating.
I also made a couple of minor tweaks to a conference website for Matthew Creasy and gave some advice to Ewan Hannaford about adding people to a mailing list. My updates to the DSL also went live this week on the DSL’s test server, and I emailed the team a detailed report of the changes, highlighting points for discussion. I’m sure I’ll need to make a number of changes to the features I’ve developed over the past few weeks once the team have had a chance to test things out. We’ll see what they say once they get back to me.
I was also contacted this week by Eleanor Lawson with a long list of changes she wanted me to make to the two Speech Star websites. Many of these were minor tweaks to text, but there were some larger issues too. I needed to update the way sound filters appear on the website in order to group different sounds together and to ensure the sounds always appear in the correct order. This was a pretty tricky thing to accomplish as the filters are automatically generated and vary depending on what other filter options the user has selected. It took a while to get working, but I got there in the end, thankfully. Eleanor had also sent me a new set of videos that needed to be added to the Edinburgh MRI Modelled Speech Corpus. These were chunks of some of the existing videos as a decision had been made that splitting them up would be more useful for users. I therefore had to process the videos and add all of the required data for them to the database. All is looking good now, though.
Next week I’ll be participating in the UCU strike action on Monday and Tuesday so it’s going to be a short week for me.
Week Beginning 4th September 2023
I continued with the new developments for the Dictionaries of the Scots Language for most of this week, focussing primarily on implementing the sparklines for dates of attestation. I decided to use the same JavaScript library as I used for the Historical Thesaurus (https://omnipotent.net/jquery.sparkline) to produce a mini bar chart for the date range, with either a 1 when a date is present or a zero where a date is not present. In order to create the ranges for an entry all of the citations that have a date for the entry are returned in date order. For SND the sparkline range is 1700 to 2000 and for DOST the range is 1050 to 1700. Any citations with dates beyond this are given a date of the start or end as applicable. Each year in the range is created with a zero assigned by default and then my script iterates through the citations to figure out which of the years needs to be assigned a 1, taking into consideration citations that have a date range in addition to ones that have a single year. After that my script iterates through the years to generate blocks of 1 values where individual 1s are found 25 years or less from each other, as I’d agreed with the team, in order to make continuous periods of usage. My script also generates a textual representation of the blocks and individual years that is then used as a tooltip for the sparkline.
I’d originally intended each year in the range to then appear as a bar in the sparkline, with no gaps between the bars in order to make larger blocks, but the bar chart sparkline that the library offers has a minimum bar width of 1 pixel. As the DOST period is 650 years this meant the sparkline would be 650 pixels wide. The screenshot below shows how this would have looked (note that in this and the following two screenshots the data represented in the sparklines is test data and doesn’t correspond to the individual entries):
I then tried grouping the individual years into bars representing five years instead. If a 1 was present in a five-year period then the value for that five year block was given a 1, otherwise it was given a 0. As you can see in the following screenshot, this worked pretty well, giving the same overall view of the data but in a smaller space. However, the sparklines were still a bit too long. I also added in the first attested date for the entry to the left of the sparkline here, as was specified in the requirements document:
As a further experiment I grouped the individual years into bars representing a decade, and again if a year in that decade featured a 1 the decade was assigned a 1, otherwise it was assigned a 0. This resulted in a sparkline that I reckon is about the right size, as you can see in the screenshot below:
With this in place I then updated the Solr indexes for entries and quotations to add in fields for the sparkline data and the sparkline tooltip text. I then updated my scripts that generated entry and quotation data for Solr to incorporate the code for generating the sparklines, first creating blocks of attestation where individual citation dates were separated by 25 years or less and then further grouping the data into decades. It took some time to get this working just right. For example, on my first attempt when encountering individual years the textual version was outputting a range with the start and end year the same (e.g. 1710-1710) when it should have just outputted a single year. But after a few iterations the data outputted successfully and I imported the new data into Solr.
With the sparkline data in Solr I then needed to update the API to retrieve the data alongside other data types and after that I could work with the data in the front-end, populating the sparklines for each result with the data for each entry and adding in the textual representation as a tooltip. Having previously worked with a DOST entry as a sample, I realised at this point that as the SND period is much shorter (300 years as opposed to 650) the SND sparklines would be a lot shorter (30 pixels as opposed to 65). Thankfully the sparkline library allows you to specify the width of the bars as each sparkline is generated and I set the width of SND bars to two pixels as opposed to the one pixel for DOST, making the SND sparklines a comparable 600 pixels wide. It does mean that the visualisation of the SND data is not exactly the same as for DOST (e.g. an individual year is represented as 2 pixels as opposed to 1) but I think the overall picture given is comparable and I don’t think this is a problem – we are just giving an overall impression of periods of attestation after all. The screenshot below shows the search results with the sparklines working with actual data, and also demonstrates a tooltip that displays the actual periods of attestation:
At this point I spotted another couple of quirks that needed to be dealt with. Firstly, we have some entries that don’t feature any citations that include dates. These understandably displayed a blank sparkline. In such cases I have updated the tooltip text to display ‘No dates of attestation currently available’. Secondly, there is a bug in the sparkline library that means an empty sparkline is displayed if all data values are identical. Having spotted this I updated my code to ensure a full block of colour was displayed in the sparkline instead of white.
With the sparklines in the search results now working I then moved onto the display of sparklines in the entry page. I wasn’t entirely sure where was the best place to put the sparkline so for now I’ve added it to the ‘About this entry’ section. I’ve also added in the dates of attestation to this section too. This is a simplified version showing the start and end dates. I’ve used ‘to’ to separate the start and end date rather than a dash because both the start and end dates can in themselves be ranges. This is because here I’m using the display version of the first date of the earliest citation and the last date of the latest citation (or first date if there is no last date). Note that this includes prefixes and representations such as ’15..’. The sparkline tooltip uses the raw years only. You can see an entry with the new dates and sparkline below:
The design of the sparklines isn’t finalised yet and we may choose to display them differently. For example, we don’t need to use the purple I’ve chosen and we could have rounded ends. The following screenshot shows the sparklines with the blue from the site header as a bar colour and rounded ends. This looks quite pleasing, but rounded ends do make it a little more difficult to see the data at the ends of the sparkline. See for example DOST ‘scunner n.’ where the two lines at the very right of the sparkline are a bit hard to see.
I also managed to complete the final task in this block of work for the DSL, which was to add in links to the search results to download the data as a CSV. The API already has facilities to output data as a CSV, but I needed to tweak this a bit to ensure the data was exported as we needed it. Fields that were arrays were not displaying properly and certain fields needed to be supressed. For other sites I’ve developed I was able to link directly to the API’s CSV output from the front-end but the DSL’s API is not publicly accessible so I had to do things a littler differently here. Instead pressing on the ‘download’ link fires an AJAX call to a PHP script that passes the query string to the API without exposing the URL of the API, then takes the CSV data and presents it as a downloadable file. This took a bit of time to sort out as the API was in itself offering the CSV as a downloadable file and this wasn’t working when being passed to another script. Instead I had to set the API to output the CSV data on screen, meaning the scripts called via AJAX could then grab this data and process it.
With all of this working I put in a Helpdesk request to get the Solr instances set up and populated on the server and I then copied all of the updated files to the DSL’s test instance. As of Friday the new Solr indexes don’t seem to be working but hopefully early next week everything will be operational. I’ll then just need to tweak the search strings of the headword search so that the new Solr headword search matches the existing search.
Also this week I had a chat with Thomas Clancy about the development of the front-end for the Iona place-names project. About a year ago I wrote a specification for the front-end but never heard anything further about it, but it looks like development will be starting soon. I also had a chat with Jennifer Smith about the data for the Speak For Yersel spin-off projects and it looks like this will be coming together in the next few weeks too. We also discussed another project that may use the data from SCOSYA and I might have some involvement in this.
Other than that I spent a bit of time on the Anglo-Norman Dictionary, creating a CSS file to style the entry XML in the Oxygen XML editor’s ‘Author’ view. The team are intending to use this view to collaborate on the entries and previously we hadn’t created any styles for it. I had to generate styles that replicated the look of the online dictionary as much as possible, which took some time to get right. I’m pretty happy with the end result, though, which you can see in the following screenshot:
Week Beginning 1st February 2021
I had two Zoom calls this week, the first on Wednesday with Kirsteen McCue to discuss a new, small project to publish a selection of musical settings to Burns poems and the second on Friday with Joanna Kopaczyk and her RA on the Scots Language Policy project to give a tutorial on how to use WordPress.
The majority of my week was divided between the Anglo-Norman Dictionary, the Dictionary of the Scots Language and the Place-names of Iona projects. For the AND I made a few tweaks to the static content of the site and migrated some more blog posts across to the new site (these are not live yet). I also added commentaries to more than 260 entries, which took some time to test. I also worked on the DTD file that the editors reference from their XML editing software to ensure that all of the elements and attributes found within commentaries are ‘allowed’ in the XML. Without doing this it was possible to add the tags in, but this would give errors in the editing software. I also batch updated all of the entries on the site to reference the new DTD and exported all of the files, zipped them up and sent them to the editors so they can work on them as required. I also began to think about migrating the TextBase from the old site to the new one, and managed to source the XML files that comprise this system. It looks like it may be quite tricky to work with these as there are more than 70 book-length XML files to deal with and so far I have not managed to locate the XSLT that was originally used to process these files.
For the DSL I completed work on the new bibliography search pages that use the new ‘V4’ data. These pages allow the authors and titles of bibliographical items to be searched, results to be viewed and individual items to be displayed. I also made some minor tweaks to the live site and had a discussion with Ann Fergusson about transferring the project’s data to the people who have set up a new editing interface for them, something I’m hoping to be able to tackle next week.
For the Place-names of Iona project I had a discussion about implementing a new ‘work of the month’ feature and spent quite a bit of time investigating using 10-digit OS grid references in the project’s CMS. The team need to use up to 10-digit grid references to get 1m accuracy for individual monuments, but the library I use in the CMS to automatically generate latitude and longitude from the supplied grid reference will only work with a 6-digit NGR. The automatically generated latitude and longitude are then automatically passed to Google Maps to ascertain the altitude of the location and all of this information is stored in the database whenever a new place-name record is created or an existing record is edited.
As the library currently in use will only accept 6-digit NGRs I had to do a bit of research into alternative libraries, and I managed to find one that can accept NGRs of 2,4,6,8 or 10 digits. Information about the library, including text boxes where you can enter an NGR and see the results can be found here: http://www.movable-type.co.uk/scripts/latlong-os-gridref.html along with an awful lot of description about the calculations and some pretty scary looking formulae.
The library is written in JavaScript, which runs in the client’s browser, whereas the previous library was written in PHP, which runs on the server. This means I needed to change the way the CMS works – previously you’d enter an NGR and then when the form was submitted to the server the PHP library would generate the latitude and longitude whereas now the latitude and longitude need to be generated in the browser as soon as the NGR is entered into the textbox, and two further textboxes for latitude and longitude will appear in the form and will then be automatically populated with the results.
This does mean the person filling out the form can see the generated latitude and longitude and also tweak it if required before submitting the form, which is a potentially useful thing. I may even be able to add a Google Map to the form so you can see (and possibly tweak) the point before submitting the form, but I’ll need to look into this further. I also still need to work on the format of the latitude and longitude as the new library generates them with a compass point (e.g. 6.420848° W) and we need to store them as a purely decimal value (e.g. -6.420848) with ‘W’ and ‘S’ figures being negatives.
However, whilst researching this I discovered a potentially worrying thing that needs discussion with the wider team. The way the Ordnance Survey generates latitude and longitude from their grid references was changed in 2014. Information about this can be found in the page linked to above in the ‘Latitude/longitudes require a datum’ section. Previously the OS used ‘OSGB-36’ to generate latitude and longitude, but in 2014 this was changed to ‘WGS84’, which is used by GPS systems. The difference in the latitude / longitude figures generated by the two systems is about 100 metres, which is quite a lot if you’re intending to pinpoint individual monuments.
The new library has facilities to generate latitude and longitude using either the new or old systems, but defaults to the new system. I’ve checked the output of the library we currently use and it uses the old ‘OSGB-36’ system. This means all of the place-names in the system so far (and all those for the previous projects) have latitudes and longitudes generated using the now obsolete (since 2014) system. To give an example of the difference, the place-name A’ Mhachair in the CMS has this location: https://www.google.com/maps/place/56%C2%B019’33.2%22N+6%C2%B025’11.4%22W/@56.3258889,-6.422022,582m/data=!3m2!1e3!4b1!4m5!3m4!1s0x0:0x0!8m2!3d56.325885!4d-6.419828 and with the newer ‘WGS84’ system it would have this location: https://www.google.com/maps/place/56%C2%B019’32.7%22N+6%C2%B025’15.1%22W/@56.325744,-6.4230367,582m/data=!3m2!1e3!4b1!4m5!3m4!1s0x0:0x0!8m2!3d56.325744!4d-6.420848
So what we need to decide before I replace the old library with the new one in the CMS is whether we switch to using ‘WGS84’ or we keep using ‘OSGB-36’. As I say, this will need further discussion before I implement any changes.
Also this week I responded to a query from Cris Sarg of the Medical Humanities Network project, spoke to Fraser Dallachy about future updates to the HT’s data from the OED, made some tweaks to the structure of the SCOSYA website for Jennifer Smith, added a plugin to the Editing Burns site for Craig Lamont and had a chat with the Books and Borrowing people about cleaning the authors data, importing the Craigston data and how to deal with a lot of borrowers that were excluded from the Selkirk data that I previously imported.
Next week I’ll be on holiday from Monday to Wednesday to cover the school half term.
Week Beginning 25 January 2021
I headed into the University for the first time this year on Wednesday this week to collect a new iPad that I’d ordered and to get some files from my office. It was great to see the old place again, but it did take quite a chunk out of my day to travel there and back, especially as I’m still home-schooling either a morning or an afternoon each day at the moment too.
As with last week, I mainly divided my time this week between the Dictionary of the Scots Language, the Anglo-Norman Dictionary and the Books and Borrowing project, with a few other bits and bobs added in as well. For the DSL I retrieved the source code for my original Scots School Dictionary app from my office so we can host this somewhere on the DSL website. This is because the DSL have commissioned someone else to make a new School Dictionary app, which launched this week, but doesn’t include an ‘English to Scots’ feature as the old app does, so we’re going to make the old app available as a website for those people who miss the feature. I also made a few minor tweaks to the main DSL site, and then focussed on adding bibliography search facilities to the new version of the API, a task that I’d begun last week.
I created a new table for the bibliographical data that includes the various fields used for DOST (note, author, editor, date, longtitle etc) and a field for the XML data used for SND. I then created two further tables for searching, one that contains every author and editor name for each item (for DOST there may be different names in the author, editor, longauthor and longeditor fields while for SND there may be any number of <author> tags) and the other containing every title for each item (DOST may have different text in title and longtitle while SND items can have any number of <title> tags). These tables allow you to search for any variant author, editor or title and find the item.
I also created two additional fields in the bibliography table that contain the ‘display author’ and ‘display title’. These are the forms that get displayed in the search results before you click on an item to open the full bibliographical entry. I then updated the V4 API to add in facilities to search and retrieve the bibliographies. I didn’t have the time to connect to this API and to implement the search on the Sienna test site, which is something I hope to do next week, but the logic behind the search and display of bibliographies is all there. There is a predictive search that will be used to generate the autocomplete list, similar to how the live site currently works: You will be able to select whether your search is for authors, titles or both and when you start typing in some text a list of matching items will appear, e.g. typing in ‘ham’ for authors in both dictionaries will display the following all items containing ‘ham’ and when you select an item this will then perform a search for the specific text. You will then be able to click on an item to view the full bibliography. This is a bit different to how the live site currently works, as with these if you enter ‘ham’ and select (for example) ‘Hamilton, J,’ from the autocomplete list you are taken directly to a page that lists all of the items for the author. However, we can’t do that any more as we no longer have unique identifiers that group bibliographical items by author. I may be able to do something similar with the page that comes up when you select an author, but this would have to rely on the name to group items together and a name may not be unique.
For the AND I made some tweaks to the website, such as adding a link to the search page if you type some text into the ‘jump to entry’ option and no matching entries are found. I then spent the rest of my time continuing to develop the new content management system, specifically the pages for managing source texts. I finished work on this, adding in facilities to add, edit, browse and delete source texts from the database. I then migrated the DTD to the new site, which is referenced by the editors’ XML editor when they work on the entry XML files. The DTD on the old server referenced several lists of things that are then used to populate drop-down lists of options in the XML editor. I migrated these too, making them dynamically generated from the underlying database rather than statis lists, meaning when (for example) new source texts are added to the CMS these will automatically become available when using the XML editor.
For the Books and Borrowing project I participated in the project’s Zoom call on Monday to discuss the project’s CMS and how to amalgamate the various duplicate author records that resulted from data uploads from different libraries. After the call I made some required changes to the CMS, such as making the editor’s notes fields visible by default again, and worked on the duplicate authors matching script to add in further outputs when comparing the author names with Levenshtein ratings of 1 and 2. I also reviewed some content that was sent to us from another library.
Also this week I responded to an email from James Caudle in Scottish Literature about a potential project he’s setting up, made a couple of changes to the Scots Language Policy website, made some tweaks to the menu structure for the Scots Syntax Atlas project and gave some advice to a post-grad student who had contacted me about setting up a corpus.






