Category: Staff Advice
Week Beginning 25th August 2026
This week I continued with the interactive map for the HiMuJe Malabar project and made some good progress with its development. All subtypes now have an associated icon that is used on the map when the ‘type/subtype’ categorisation is selected, and I also reworked the colours of these, as you can see in the following screenshot:
After last week’s frustrations with the map legend, I reworked it to divide it a little better into sections by type, which you can also see in the above screenshot. When a location popup is opened the corresponding subtype icon also now appears in the popup header. I also added an option to the ‘Home’ menu to make all marker labels permanently visible on the map (until you turn them off again). With all locations on this does clutter up the map, but I think it’s a useful option for people who are looking for a place but aren’t entirely sure where it is. With labels on you can see the names of all places without having to hover over each marker in turn. This works better when the map is filtered in the legend. I also intentionally set the map so that I you set the labels to on and open the travel route menu the labels are removed as the travel routes would be very difficult to see with the labels on.
I also updated the colours for the other categorisation types to hopefully make them work a little better. When the ‘number of references’ categorisation is selected the marker colours are now a gradient. It’s also now possible to cite a specific stage in a travel route. This appears in the citation text, e.g. “HiMuJe Malabar Digital Map: Watercolor map, travel routes menu, showing the travel route ‘Brahmins’, with unrelated places hidden, viewing stage 2 of the route, placenames categorised by language. 2026. In HiMuJe Malabar. Glasgow: University of Glasgow. Retrieved 28 August 2026 and following URL opens the route at Stage 2. This option will mean I’ll be able to update the location popup to add a new tab that lists all of the stages in routes that the place appears in, and give links to each of these. I haven’t done this yet, though.
Also this week I spent quite some time working for the Dictionaries of the Scots Language. This included attending an online meeting to discuss the structure of the new site, at which some fairly major revisions to the new design were proposed. I also engaged in a couple of email conversations about the updates that still need to be undertaken before the hopefully launch the new site later this year. In preparation for all of this work I also went back through all of my emails from the DSL team and my notes to create a bit ‘to do’ list containing everything that needs done. It’s going to be a busy couple of months.
Also this week I met with Craig Lamont and his RA Emily Hay to discuss the Life Writing project. They’re going to be compiling a bibliography of Scottish life writing and I’m going to develop the systems for this. I did look into using existing bibliography software but nothing I looked at did exactly what they are after so I’m going to create a simple CMS for them to use myself, with a front-end consisting of search and browse options and a map interface sometime next year. I spent a bit of time after the meeting going through the sample spreadsheet they had prepared and sending a list of questions about the data.
I also met with Deven Parker this week to discuss an AHRC proposal she’s involved with. I can’t say too much about this, but it’s related to the work we’re doing on the playbills and I would be involved in some technical capacity. I also had a chat with Garrick Allen about a proposal he’s trying to get funding for after an earlier AHRC submission was unsuccessful. I read through the documentation and gave feedback, and I’ve also put Garrick in touch with my colleague Luca Guariento who is probably going to provide the technical assistance.
Other tasks I carried out this week included investigating an issue with the Anglo-Norman Dictionary (which turned out to be an issue with the data and not the system), making some final tweaks to the Burns Supper Map before the project RA finished working on the project, and generating some Books and Borrowing data for a publication Katie Halsey is working on.
Week Beginning 3rd August 2026
I worked a total of four days over the past two weeks, and was on holiday for the remainder. During this time I had a meeting to further discuss a place-names related AHRC proposal I’m involved with. I can’t really say much more about it at this stage, but the proposal is coming together. I also had to spend some time working with IT Support and Luca to figure out why our local server kept going offline repeatedly. It looks like this was caused by the server getting swamped by requests from one particular source (almost certainly bot or AI) and thankfully IT Support were able to block this, after which the server was stable again. It’s something we’re going to have to keep looking out for in future. I also spent a bit of time working with Luca to get automatic WordPress updates working on our local server, as the way sites had been set up meant that the setting was not working. Luca managed to find a solution to this, which is really great.
I also spent a bit more time working on the Burns Supper Map, creating a record for it on this site (see https://digital-humanities.glasgow.ac.uk/project/?id=156), adding more suppers that had been submitted via the survey and making some requested edits to existing suppers. I also set up access to Google Analytics for the other two members of the project team.
In addition, I investigated an issue with the Scots Syntax Atlas after someone suggested that the linguists’ atlas was looking somewhat blurry. I managed to figure out why this might be the case, although I’m not entirely sure whether this is a new issue or if the markers always looked like that. I contacted the project PI and suggested a couple of updates, but I haven’t heard back yet so will need to wait and see what she says.
I spent most of the remainder of my time working on updates to our new test interface for Dictionaries of the Scots Language and working through the list of outstanding items for the DOST Auld Laws project. For the DSL I completed the updates to the bibliography page that I began working on a couple of weeks ago. I implemented pagination of the entries associated with bibliographical items, with navigation bars appearing above and below the entries, with 20 appearing per page and ‘jump to page’ buttons also appearing, just like with the search results. This works pretty well, but is somewhat cumbersome for someone like Sir Walter Scott, who is referenced in 2861 entries, split over 144 pages. We don’t have this issue with the search results are these are capped at 500 (25 pages) so we might need to think of other ways of handling this.
I also ensured that headword searches that don’t yield any results automatically perform a fulltext search for the term supplied. This works on the live site, but only when both dictionaries are selected. With the new site there are separate quick searches for SND and DOST so the additional search wasn’t being triggered. It is now, as is the advanced headword search for both dictionaries.
I also made tweaks to the DSL’s new regional map based on feedback I’d received – adding in some content where we previously had placeholder text and ensuring the ‘About’ popup didn’t disappear off the bottom of smaller screens and a few other small updates. I then began to look at the ancillary pages and how we can make them look a bit nicer. I spent a bit of time on the ‘Word of the Week’ page and liaised with William Ashford, who is responsible for such content about this, and further updates that were going to make to the ancillary content closer to the launch date of the new site (which will hopefully be in November).
For the DOST Auld Laws project I added the top navigation bar that will link the site in with the SCOTS Corpus and CMSW. I added in copyright information and added facilities to download page images and the XML files for each document. These being up a pop-up asking for users to abide by the license before leading to the actual content, which hopefully won’t be too annoying. I also added in the ‘cite’ popup to all document pages, which took a little time to implement, and added in Google Analytics. I also made the image thumbnails on the document overview pages smaller and placed them in a collapsible section that is closed by default, plus I removed the introduction to the documents page, as this will be covered by the homepage.
I also removed some pages that didn’t have content (e.g. blank pages) from the beginning and end of some of the documents and I added a feature to turn off and on the line highlighting feature. The highlighting feature allows the user to click on a line of text in the image or text for a page and for that line to be highlighted in both the image and the text, which is pretty nice. Unfortunately the line highlighting gets in the way of the image viewer’s zoom and pan functionality on touchscreens, making it a somewhat unreliable and frustrating experience. This new feature removes the option to ‘click’ on a line, meaning pointer events are not intercepted and make their way reliably through to the image viewer, which works much more smoothly.
Also this week I had an email conversation about user feedback and walkthough videos for the STAR resources and booked my accommodation for the DHC conference in Sheffield. Next week I’m back in Glasgow and back working a full week, with summer holidays all over.
Week Beginning 4th May 2026
This was a four-day week as Monday was a bank holiday. I spent most of Tuesday and Wednesday continuing to work on the DOST Auld Laws project, working with the XML files. The XML files generated by Transkribus and exported by the tool as TEI XML contained many elements that were not valid TEI elements, such as the custom <Aitken> element that had been applied to notes added by A J Aitken. My first task of the week was to write and apply transformations to the XML files to convert them into fully valid TEI. I achieved this using XSLT, which is a language I find very unintuitive and frustrating to work with, no doubt exacerbated by the fact that I don’t work with it very often. I struggled to get any transformations to run initially, and ended up turning to AI to figure out why my scripts were not working. In this instance AI proved to be extremely useful as it identified what the problems were (e.g. I hadn’t declared the correct namespace or used it when writing the rules) and really helped me to understand how everything fitted together. I still wrote the bulk of the code myself, but AI was very helpful in identifying errors or issues. By the end of Tuesday I had generated (and checked) a collection of TEI files that successfully validated in Oxygen, which was a good milestone to reach.
On Wednesday I then worked on the front-end for the project, figuring out how to transform the valid TEI XML into HTML for display in the ‘text and image’ and ‘text only’ views of document pages. As this was once more using XSLT I enlisted the help of AI to figure out specific issues that I was unfamiliar with. The biggest of these was how to pick out and process one single page from a document’s XML file based on the <pb/> element. I had no idea how to achieve this, and despite this being a fairly fundamental issue when processing TEI documents I didn’t manage to find any useful information online. However, AI came up with a solution (and equally importantly an explanation) in a few seconds and I was then able to incorporate this into my code, transforming the contents of one page, whose ID was passed to the XSLT script as a parameter, to HTML with a variety of styles applied to the various elements.
I also updated the ‘click on a line in the image’ feature, and it’s now possible to deselect the line if it’s already selected – previously once you’d highlighted a line you couldn’t get rid of the highlighting, only move it to a different line but now if you press on the highlight it’s removed. The lines are also now connected to the text view – pressing on a line in the image also highlights the corresponding line in the text. You can also press on a line in the text to highlight it and the corresponding line in the image.
I also updated the height of the text pane so that it matches the height of the image pane and if the text is longer the pane scrolls. This ensures that if the text is very long it’s still possible to see the image, rather than having the entire page scrolling, which may result in the image not being visible when you’re at the end of the text. It is how we did things in Books and Borrowing, but having a scrollbar in a section of the page in addition to the browser’s scrollbar can be annoying for some people so I might revert to the previous layout depending on feedback. Here’s a screenshot showing the transformed text, a highlighted line and some of the formatting that’s been added:
Also this week I continued to work on the Place-names of Armagh project. I wrote and executed a script that generated Irish grid references and ITM values for all places based on their latitude and longitude, a task that I completed using the ‘GridRefUtils’ scripts as detailed here: https://www.howtocreate.co.uk/php/gridrefapi.php. I’ve used these scripts before on previous place-name projects and they’ve been hugely useful. I then ran a further script to generated altitude for the place-names by connecting to the Google Maps API, so we now have complete geospatial data for all of the places (other than the 386 that didn’t include Easting and Northing data in the original spreadsheet). I also added in a new parish and barony in a different county that one place-name requires, and had discussions with the team about the splitting of analysis data into Irish forms and translations. Removing the Irish forms from the translation field is going to be rather tricky to automate as the Irish forms often form part of the translation. It’s looking like I’ll need to automate the transformation of some of these, with the rest then requiring manual intervention.
I also continued working on the redevelopment of the Mapping Metaphor resource to create a combined English and Old English map, something I began last week. This week I managed to get the combined view of the drilldown of the visualisation working. The counts represented by the yellow circles also use the combined data, and these are also displayed in the pop-up card, for example, if you select 1K02 Creation and press on the yellow line or circle for 1B Life. In the combined map there are connections to 6 categories in Life whereas there are 5 in the E map and 2 in the OE map (one of the OE ones is also present in the E map, which is why the combined total is 6). The combined pop-up card also now lists the number of OE lexemes in addition to the E lexemes, for example “1K02 Creation 1088 lexemes / 127 OE lexemes”.
I still need to implement the combined visualisation view of the search results, and also the card pop-up between individual categories, which is going to take quite some reworking (possible direction changes, combined example lexemes, timeline updates etc). I’ll hopefully find some time to continue with this next week.
Also this week I processed and added some more images to a few Burns Suppers, had a chat with Garrick Allen about a research project he’s wanting me to be involved with later in the year, and contacted Lindsay Balfour about a proposal she’s writing that will have some technical requirements. I also made a couple of tweaks to the Thesaurus of Old English website after Fraser go in touch with some suggestions, made some tweaks to the survey popup on the Dictionaries of the Scots Language website, gave some mapping advice to Renu of the HiMuJe Malabar project, and created an alternative version of one of the Anglo-Norman Dictionary’s textbase documents that strips out all Latin text.
Week Beginning 30th March 2026
This was a three-day week for me because Friday was Good Friday and I’d taken the Thursday off as well. Despite this, I still managed to work on four different projects. Last week for the Playbills project I’d written a script to process performer names, including splitting this up into titles, forenames, surnames and other names, and also assigning gender. What I’d spotted and hadn’t had time to address is that there are several thousand performer names that are actually multiple performers that would need to be split into individual people. I spent most of Monday working on a script that would split these multiple performers up and process them all individually. This managed to reduce the number of unprocessed performers from 1271 to 414 (and most of the ones remaining are either blank or have text like ‘performer unknown’). I saved all performers in spreadsheets so Deven can check them – as of yet I haven’t made any updates to the database. There are 48,234 individual performers that were successfully processed by my script and a further 3,491 that were processed from multiple performers, leaving just 414 unprocessed performers, so I’m pretty happy at how successful my scripts have been.
Also for the Playbills project this week I had further discussions with Deven about downgrading certain plays to ‘special attractions’, which are things like songs and dance that are not full plays. Deven had also completed work on the spreadsheet that mapped out which performances involved the same plays (even if their titles were not exactly the same) that I will use to generate canonical records for plays. This spreadsheet also notes which plays should be downgraded and others that should be deleted entirely.
Whilst working on the scripts that will generate canonical records, downgrade plays and delete others I spotted some issues with the spreadsheet, such as plays being marked for deletion that looks like legitimate plays to me. Some other rows in the spreadsheet had been marked as both to be merged and deleted, which wasn’t right. I sent a list of possible issues to Deven and I’ll need to take this up with her once I’m back from my holiday.
For the Burns Supper Map project I had an in-person meeting the Cleo and Pauline to discuss the map, the interface, the data and our plans for the weeks ahead. It was a great meeting and really useful to meet up in person. Everyone is very happy with how the map interface is coming along and we have decided which of the test interfaces I’d developed would be used for the live site (mostly the first test interface but with the fonts from the second one). I have a 14 point ‘to do’ list for the project that I’ll work through once I’m back from my holiday.
For the HiMuJe Malabar project I wrote a non-technical specification document for the map I’ll be developing for the project, based on the discussions I had with Renu and Ophira at our meeting last week. I’ve got a pretty good idea about what needs to be developed now and I sent the first draft of my document to Renu and Ophira for feedback.
The remainder of my time was spent trying to digitise an old cassette tape of Scots and Gaelic poetry for Alan Riach in Scottish Literature. I had hoped I’d be able to plug an old tape player into my laptop to do this, but then I realised my laptop (and most modern computing equipment) doesn’t have a line-in port so this wouldn’t work. I then asked Jane Stuart-Smith whether there may be facilities in the GULP lab that I could use, and while there was a tape player we were unable to get it connected to any computer in the lab. I finally managed to digitise the tape by digging my old Hi-Fi out of my attic at home and plugging it into an old PC. Thankfully this worked and I recorded each individual poem as a separate MP3, then sent these to Alan.
I was off on Thursday and Friday, and I’ll be off for Easter for all of the following week and will return to work on Monday the 13th of April. So that’s all from me for now.
Week Beginning 23rd February 2026
I continued to work on the Burns Supper Map from Monday to Wednesday this week, with my first task being to import the public domain data into the map, taking the total number of suppers on the map to 751. I then implemented the advanced filter options. You can now open the advanced filter popup and select any of the filters you’re interested in. When you press the ‘apply filters’ button at the bottom of the popup it closes and the map displays only those suppers that match your criteria. The ‘Advanced’ section of the side menu then displays the number of matching suppers, your selected filters and buttons to refine or choose new filters, as the screenshot below demonstrates:
Note that I haven’t done any work on the colour schemes for the map yet – it’s all just using the colours taken from the Iona place-names map, but this will change. I also updated the map so that both the advanced filter and regular filter options now reposition the map to show all matching suppers when selected. However, icons on the left of the map can get obscured by the map menu. It has a tendency to sit on top of the US. I’m not sure what to do about this – I could ensure the map zooms out further, or I could close the side menu, although this might just confuse people.
I also updated the non-advanced filters section so that it displays the total number of suppers. However, a supper may have multiple filter options in a filter type so this total will not be the same figure as adding up all of the counts for individual options (e.g. one supper can have multiple toasts so will appear in the count for each individual toast that it features). I then created the table view. Pressing on the ‘Table view’ button will display all suppers currently found on the map in tabular form, and you can reorder the rows by pressing on a heading (e.g. ordering the rows by country). Pressing on the venue name link closes the table view, centres the map on the supper and opens the supper’s in-map record. Note that if you’ve performed a filter then the table only contains the filtered data.
Setting up the advanced filter option was especially time-consuming to implement so it was good to get that finished. There’s still a lot left to do, such as ensuring filter options get added to the page URL to enable bookmarking / sharing / citing of specific results. This is going to be another big job, and one that I’ll hopefully tackle next week.
I also spent some time this week drafting some text with Luca for a page about the technical developers across the College and the services we offer. This is not yet live, but will be a useful information point for staff who are looking to create an online resource for their data.
On Tuesday this week I participated in a meeting to discuss a new proposal being led by Sara Pons-Sanz at Cardiff that Glasgow will be involved with. It was a useful meeting and we all have a clearer idea about what the project will entail and how Glasgow will contribute. I had a further meeting on Friday with Jennifer Smith and Janine Illian about statistical analysis of the Speak For Yersel data and it looks like we’ll be getting some people in statistics working with our data, which is great.
I spent the remainder of the week working on the Playbills project. I’ve begun to set up an API and pages that will allow people to browse the playbills data. We still have a lot of work to do with the data, such as creating single, canonical records for venues, plays and other data types, so it’s likely that all of the data will need to be replaced at some point, but as the data structures are mostly finalised I decided to start work on some aspects of the front-end. So far I’ve created pages for browsing venues, listing playbills and browsing genres. I hope to continue with this next week.
Week Beginning 16th February 2026
On Monday this week I had a meeting with Joanna Kopaczyk and Pia Lehecka to discuss the DOST source materials transcription project. The materials have all been transcribed using Transkribus and the project needs a front-end created through which the images and text can be searched and browsed. A mockup interface has already been very kindly produced for us by Dario Kampkaspar, who created a pathway for converting Transkribus XML to TEI XML (see ‘Page2tei’ here: https://help.transkribus.org/downloading) so hopefully it won’t be too difficult to create something similar for the entire dataset, as task I’ll need to have completed by the end of May.
On Tuesday I had a meeting with Deven Parker regarding her Playbills project. We met with a PhD student in Computing Science who has been developing the pathway for sending playbill images to AI and processing the files that are returned. We needed to make a few tweaks to this pathway, and also to figure out how we can run things ourselves and it was a really useful meeting to participate in. After the meeting Deven and I had a further meeting to discuss the data extraction and processing I’ve been doing, and subsequent data cleaning and amalgamation tasks that will need to be undertaken. Deven is going to try and proofread and correct the YAML files that were originally generated by the AI as an initial step, and this will mean I’ll need to delete my data and regenerate everything from this updated dataset. I can’t therefore do much with the data I currently have, but will instead focus on developing the methods for browsing the playbills in the front-end, something I’ll begin working on next week.
I also met with Louis Strange on Tuesday this week to discuss a new proposal he’s putting together. It was a useful meeting and I gave him some advice about possible ways his data could be presented and used online. I can’t say much more about the proposal at the moment, but hopefully it will be funded.
Other than meetings, I spent quite a bit of time this week working on the interface for the new interactive map of Burns Suppers. I set up the basic map interface (using a satellite map with labels) and set up the map menu using the components I created for the Iona place-names project. For now the side panel has the same colours as the Iona map, but this will be changed in time.
My first task was to ensure that all of the supper records could be loaded into the map, and these work with the clustering tool. Currently the map marker is the same as the old supper map (red background, knife and fork icon). I did try using a haggis icon that we found online, but it’s not really going to work as when scaled down to a suitable size it just looks like an indistinct blob. I also tried using the Font Awesome ‘lemon’ icon as it looked vaguely like a haggis, but I’m not sure this is ideal either. We’re going to need to give this some further thought.
I added in all of the textual information about the suppers to the record popup and the map is already beginning to shape up quite nicely. I then added in a facility to allow you to bookmark / share exact views of the map (position and zoom level) using the Leaflet Hash plugin. Eventually the URLs will also feature filters etc but I haven’t implemented that yet. But I have implemented the filter options themselves, all of which should now be working. If you press on the ‘Filters’ menu item you should now see some text about the filters and a drop-down list featuring the various filter types. I realised my specification document hadn’t included number of guests as a filter type so I’ve added that in too. One you select a type (e.g. ‘Poems and songs’) the individual filter options are listed, together with counts of the number of suppers that feature them. You can then press on an option and the map will update to only display the matching suppers. The screenshot below shows the map with ‘Tam o’ Shanter’ selected. Note that as of yet the map doesn’t zoom and pan to ensure all matching suppers appear in view – instead the map just stays where you previously had it. I’m not sure now whether making the map automatically change position is a good idea or not, so it’s something to think about.
I also updated the ‘Home’ menu to add in the title of the resource and to make the buttons work. The ‘help’ button now opens a popup with some placeholder text while the ‘reset’ button resets the map to include all data and the default position and zoom. I’ve updated this as previously it was zoomed in on Scotland. Now the map is zoomed out and shows much of the world (depending on your screen size). I also ensured the ‘Attribution and copyright’ link in the bottom right now works, displaying a popup with some placeholder text, and I’ve added some text to the ‘Advanced’ menu section, although the advanced filter options are still to do. I also tested all this on my phone and everything currently in place works fine on it. Next week I’ll probably work on the advanced filter options, which will allow users to combine different filters.
I spent most of Thursday updating the APIs for various sites, as Luca had spotted some inefficiencies that could be improved. I’ve now updated all of the APIs that we host locally, but I still need to implement the changes on our externally hosted sites.
On Friday I worked for the Dictionaries of the Scots Language, investigating some non-urgent issues with searching that had been sent to me in November last year. It turns out that these issues are all sorted with the new DSL interface we’re hoping to launch this year, so I didn’t need to make any changes to anything. I also made some updates to the layout of some aspects of the new tags on the new DSL website, and the rest of my time was spent making further tweaks to the new DSL website that have been on my ‘to do’ list for a while.
I’d spotted that the search results in the side panel on entries was including blank links that would be highlighted when hovered over and I managed to identify the issue and sort this now. It turns out I’d already sorted this for searches that involve both SND and DOST, but hadn’t applied the update to individual dictionary searches. I also created a new version of the bibliography page that lists the actual quotations as well as providing links to the entry pages. I can’t really share any screenshots of this for now, and I don’t know if we want to include the quotations or not, but I thought I’d create this test version so we can see how it might work.
I also made the links from the bibliographies on both the test version and the main version take you to the actual quotation within the entry, rather than dumping you at the top of the entry. I used our purpley-pink highlight colour to highlight the citation to make it easy to spot which one brought you to the entry. This actually took quite some time to implement as a bibliographical item can be cited multiple times in an entry and in such cases the system needs to know which one to link to, but I reckon linking through to the actual quotation from the bibliograph page will be a hugely useful addition.
Week Beginning 2nd February 2026
I had a bit of a disrupted week this week, as I started feeling unwell on Tuesday morning and ended up off work sick for the rest of Tuesday and Wednesday. During this time I felt absolutely wiped out and could barely do anything other than sleep, but by Wednesday evening this had developed into a monstrous cold, the likes of which I’ve not had for several years. Thankfully once the symptoms had moved to my nose and throat my head was a bit clearer and I was able to work on Thursday and Friday, but I was still pretty far from feeling 100%.
I spent most of Monday this week preparing for, travelling to and co-presenting a talk about Speak For Yersel at the Edinburgh Futures Institute with Jennifer Smith. The talk went pretty well and it was good to meet some of our linguistics colleagues at Edinburgh, plus others involved with the EFI. I spent some of my other available time reading through and commenting on an AHRC proposal that will involve Glasgow and the Historical Thesaurus that had been sent by Sara Pons-Sanz at Cardiff University, and looking through some further place-name data I’d been sent for the Place-names of Armagh project.
Despite being off work sick on Wednesday I still managed to attend an online meeting for the Burns Supper Map project to discuss the specification document I’d prepared for the project. This was all very positive and there weren’t any major issues that anyone had spotted whilst reading through it.
For the remainder of the week I spent a bit of time investigating some issues that had been encountered when publishing pure xref entries through the Anglo-Norman Dictionary’s management system. Certain cross references were not appearing in the published entries despite being in the XML and a bit of investigation uncovered why. The entries contained cross references to entries that don’t actually exist in the dictionary. For example, Mars_2 references ‘march’, which is not an entry and respundre_2 references ‘repundre’ which is also not an entry (they both need homonym numbers added). When xref entries are published the cross references are extracted and stored, and at this point the system checks that the references are valid, and only links to entries that are valid. It is these that are displayed in the front-end, so even though invalid xrefs may exist in the XML they don’t get displayed. The ‘preview’ generates its view directly from the XML without checking validity, which is why this view doesn’t match the front-end. I ran a check and it turns out that there are around 500 xref entries that include a reference to an entry that doesn’t exist, and I passed these onto the editor who will get these sorted.
On Friday I met with Jennifer and Janine Illian, who is the current Head of Statistics, to discuss the Speak For Yersel data and what kind of additional statistical analysis might be possible. Janine is particularly interested in spatial modelling and has a keen interest in linguistics and it was really great to hear her thoughts about the Speak For Yersel data. I’m going to send her the data for all survey responses next week so she can experiment with it, and we’ve arranged to meet again later this month.
I spent the rest of my available time this week working on Deven Parker’s Playbills project, working with the YAML files, figuring out how these might be imported into Solr and how we can extract canonical records for things like venues from them. It turns out that Solr can’t index YAML files (at least not without creating a custom data importer), which is a bit of a surprise. This isn’t a major issue, though, as I can convert them to JSON, although this also proved to be trickier than I’d anticipated. Normally I’d use PHP to process data, but PHP also can’t read YAML files, at least not without installing extensions and this process seemed far too convoluted to bother with. Instead I used Python to convert the files, but this involved a bit of trial and error as I’m not used to Python and it’s bizarre insistence on whitespace being important, and the fact that if you mix up spaces and tabs to create this whitespace the scripts fall over. I got there in the end, though.
The bigger issue I encountered was with the unit of data that gets indexed. I’d previously said that we’d index entire playbill files and use ‘playbill’ as the smallest item that gets returned in the search results, but it turns out there are some problems with this, and I think indexing individual plays is going to work better. I’m still experimenting with the data and Solr’s capabilities, but initial impressions are that it isn’t very good when working with subsets of data within individual files, or more complex queries. For example, you can search the playbills for the title ‘Macbeth’ and find matching playbills. But if you combine this with another field that exists in another play in the playbill (e.g. role ‘Jacques Strop’) the playbill record will still be returned. So even though the role mentioned actually belongs to a different play in the playbill, because both pieces of information exist somewhere in the playbill it gets returned.
With my initial experiments Solr also flattened out the data – all performer names appear in one list per playbill, not separate lists per play, and it’s the same with roles. Other than the order of the items in the lists, there is nothing to connect the two. The following screenshot shows one playbill record indexed within Solr (just using Solr’s default post and without customising a schema). You can maybe see how Solr has flattened things out, resulting in data being lost (e.g. which performer belongs to which play).
I then tried to index the data at play level, adding in a play ID and also any playbill level data (thus ensuring it’s still possible to search for date, venue etc). You can see the results in the following screenshot, which includes 5 separate records.
Here at least it’s possible to tell which performer / role belongs to which play. But performers / roles are still only connected by their position in the lists. Record 5 is a duplicate I made of record 4, but I deleted the ‘role’ text for one performer to see what would happen. And Solr indexed the record as it was, with 5 performers and 4 roles, so based on list order ‘Miss Newton’ is now ‘Landlord’ and not ‘Marie’, and ‘Mr. Watkins’ now had no role.
After further investigation I realised that it is possible to get sole to properly index nested data (see https://solr.apache.org/guide/solr/latest/indexing-guide/indexing-nested-documents.html) although instructions on how to actually import nested data into Solr are pretty thin on the ground – you can’t just use the default ‘post’ command as this flattens all data. I ended up following another tutorial (see https://docs.arenadata.io/en/ADH/current/how-to/solr/solr-index-nested-docs.html) and importing the data using the Solr admin interface. This thankfully worked, as the following screenshot demonstrates. You can see that individual performers are directly associated with roles.
There’s still a massive amount to do with the data, though. I need to extract unique venues, plays, performers and roles and assign IDs to them to enable them to be searches for. I decided that it would be easier to manage such processes via a relational database, so on Friday and mapped out a structure for the playbill data and bean working on an import script that would process the JSON files. Lots more to do in the coming weeks!
Week Beginning 19th January 2026
I worked on many different projects this week, but the one I spent the most time on was the Dictionaries of the Scots Language. I’ve not done much work for the DSL since the intensive period I spent developing the new website interface and deploying it on our test server ahead of the face-to-face meeting in mid-November. I had a list of further updates I needed to make following on from this meeting, but I needed to work on other projects since then and hadn’t got around to it. I’d also received a number of emails about changes to the presentation of entries reflecting the structural changes to the entry XML that I’d put to one side.
On Wednesday I had an online call scheduled with the DSL team to discuss the new front-end and it seemed like a good opportunity to get back to grips with all of my outstanding DSL tasks. This mainly involved making updates to the XSLT on our test server to tweak the layout of various items in the new entry XML structure, such as adding commas between tags when they are rendered, ensuring certain tags or attributes that weren’t getting rendered before appeared in the generated HTML, updating the styles of certain elements like the content warning labels, fixing a few bugs such as the ‘sticky’ heading not displaying in certain circumstances. There were at least 20 such items that needed investigating, fixing and testing, so this took quite some time to work through, but I managed to complete it all during the course of the week.
The meeting itself was very useful and as always it was good to catch up with some of the other DSL team members. There’s going to be a big push towards getting the new website interface ready for publication this year, and I’m obviously going to be involved in this process. I already have a number of items I need to sort out with the new interface and I’ll try and get started on these over the coming weeks.
I also spent a bit of time this week working for the Anglo-Norman Dictionary, investigating a strange occurrence with the publication of updates to entries, which turned out to be a user rather than a system issue, reinstating the links out from entries to the DMF dictionary, as their website is now properly back online again, and tweaking the wording of the quick search and ‘jump to entry’ text throughout the site.
I also did small amounts of work for several other projects, such as updating the licensing statements across the Seeing Speech and Star sites, fixing an issue with the ‘download song’ facilities on the Editing Robert Burns site, sorting an issue with the HiMuJe Malabar site, submitting my expenses from the Zurich workshop, exporting some SCOSYA data for Jennifer Smith, helping to sort out an issue with the Helsinki Corpus, and having a conversation with Clara Cohen about a new proposal she’s putting together. I also made some further updates to the VARICS look-up system, adding in some introductory text, some further references, and reworking the measurement processing so that when a red or amber result is given new textual sections about what this means and what the next steps should be appear underneath in collapsible accordion sections.
Also this week I had a meeting with the Burns Supper Map team to discuss the data that is now coming in and how and when I should start working on a new interactive map to visualise it. I also put in a request for a new subdomain for the project that was set up by the end of the week. Next week I’ll probably write a brief specification document for the front-end.
My final project of the week was the Place-names of Armagh project, for which I started working with some existing place-name data for the area. There are around 230 place-names and several thousand historical forms and I spent quite some time researching how the data was structured and how it might be mapped onto the Glasgow place-names system. This included analysing the geospatial data, including shapefiles for Townlands and what I though was Parishes (but actually turned out to be the same as for Townlands). I had hoped to be able to import the data by the end of the week, but my analysis of the data raised a lot of questions that still need to be addressed, and I’ll need to continue with this next week.
Week Beginning 17th November 2025
After spending a pretty intensive few weeks working on the new interface for the Dictionaries of the Scots Language ahead of last week’s in-person meeting, this week I was able to return to other projects that I’d had to put to one side recently. I am still unfortunately suffering from a rather bad bout of sciatica, which is now in its third week and is making it hard to work, especially in the mornings. It has also unfortunately prevented me from travelling to the University, which meant I had to rearrange a couple of in-person meetings this week. Despite all of this I’ve still managed to get quite a lot done this week.
I spent a lot of my time this week working on the migration of the place-names of Fife data, which I’d started to look into again last Friday. I hadn’t had any time to work on this since September and it was good to get back into it. I’m migrating the Fife data, which I originally extracted from a Word file into a relational database structure way back in 2016 to the same structure that I created for other place-names projects such as Berwickshire and Iona. As the Fife data is very messy there’s much work to be done to get it ready.
This week I managed to complete work on the historical forms. There are more than 23,000 historical forms in the data, and for Fife the sources of these forms are stored as a field in the historical form table, meaning there are more than 23,000 sources. For the other projects an individual source is stored in a separate ‘sources’ table only once, and then is connected to all relevant historical forms via a joining table that also stores the specific reference for the form (e.g. a page number).
What I needed to do for the Fife data was to extract the unique sources, separate out the reference data (which was stored as part of the same field), insert new sources once into the ‘sources’ table and add in references to this source for each historical form where the source appears. Thankfully there were some patterns to the source data, plus the same sources kept cropping up for many of the historical forms. For example, the source name was italicised in more than 11,500 historical form records, which made it easy to split up the source name from the reference and process these sources.
Of the remainder there were a number of major sources, such as ‘OS 6 inch 1st edn’, found in almost 2,200 historical form records. I was able to process sweeps of the data that ticked off hundreds or thousands of historical forms at a time, which then left me with the awkward records that needed more manual intervention. Even in such cases I was able to automate the process to a certain extent by first finding where the source name ended and the reference began, and then for each historical form that featured the source splitting the field at this point and storing the source and the reference. I manage to sort out all 23,000+ historical forms this week using various methods, resulting in less than 500 unique sources being stored. This is of course still just a first draft and there will almost certainly be some duplicates due to different spellings and such things, but the source data is now in the right format and is clean enough to be managed through the place-names CMS system, once I get round to setting it up for this project.
I also began work on the final data migration I’ll need to tackle for the Fife data: the place-name elements. As with the source data, these are not stored in an especially relational way. In the more modern place-names projects each element is stored once has an associated language that is stored once. When elements appear in a place-name they are then associated by means of a joining table that stores a reference to the element and the place-name, and information about how the element appears within the place-name, such as its position and how it is connected to a subsequent element. In Fife this is all just stored in one table, meaning there are almost 6,500 elements. I began cleaning this data up a bit this week, as I’d spotted some issues, such as elements being stored without a language due to the data from the original Word file not being processed successfully. I’ll continue with this next week, if I have the time.
Also this week I met with Ophira and Renu to discuss the structure of the spreadsheet that I created to store details about the places that will appear on the map for Ophira’s HiMuJe Malaber project. This should have been an in-person meeting but as I am still unable to leave the house much we had the meeting online instead. We went through the spreadsheet in detail and several structural changes were proposed. After the meeting I then spent some time making the updates and ensuring these were logged in the accompanying data dictionary.
Also this week I had a lengthy email conversation with Andrew McHugh, Luca Guariento and others about hosting the online resource for one of Garrick Allen’s projects. This is currently hosted elsewhere but Garrick would like it to come to Glasgow. There was a lot to discuss about this request as the resource uses technologies we don’t otherwise support at Glasgow, but by the end of the week we’d reached a decision about how and where the resource should be hosted, and Luca has agreed to oversee the process.
I also found a bit of time this week to finally swap the live Historical Thesaurus website with the new version I created earlier this year that uses a new, unified API, unlike the mess of scripts that were cobbled together over a decade or so of development that previously powered the site. This is a major update to the resource’s back-end but includes no changes whatsoever to the front-end, so all of the work that went into it should be invisible. But it will make it much easier to manage the resource in future as all data access now passes through one script.
I had an in-person meeting planned with Deven Parker this week to discuss her playbills project, but I was unable to travel into the University and instead we had an email conversation. Deven is just about at the stage for me to begin creating an online resource for the data, and she shared with me the YAML files that had been generated by the AI ‘reading’ the playbill images and extracting and formatting the data. My first task was to identify which of the playbills had been classified as ‘melodrama’ as it is these that Deven wants to initially focus on. Deven needed some help in creating a Python script that could export a list of matching YAML files, and this was a good opportunity for me to learn a bit of Python, as this is the kind of activity I’d normally just write a PHP script for. I managed instead to write a Python script that did was Deven needed, and we had a lengthy email conversation about the data, the project and the next steps. I’ll hopefully be well enough to meet her in person soon to define some requirements for the resource she’d like me to build soon.
Week Beginning 10th November 2025
I continued to work on the new interface for the Dictionaries of the Scots Language this week, ahead of the DSL’s face-to-face meeting on Wednesday at which I gave a demonstration of the interface. There was a lot still to do before the meeting, and I ended up working several hours over the weekend to ensure everything was ready. By Tuesday I’d managed to complete just about everything, although there are still a few formatting issues with some of the ‘history’ pages that I’ll need to tidy up at some point. With the completed interface still only running locally on my laptop the next step was to deploy it to our online test instance of the DSL site, which thankfully only took a couple of hours to complete.
I then spent most of the rest of Tuesday writing and preparing the text of a walkthrough of the new interface for Wednesday’s meeting. Of course, whilst creating this walkthrough I spotted several minor issues with the site and spent some further time addressing those too. I didn’t need to be at the meeting in Edinburgh until the afternoon so in addition to running through my demo a few times on Wednesday morning, I also added a feature to the site’s bibliography page that I’d been meaning to implement for a while. The bibliography page as it currently stands is a bit of a dead end on the site: you search for an author or title and you reach a bibliographical record for a specific work or works. There is nowhere to go from there – no pathway to take you from the record to the entries that cite the work or works. We store this information in the database already, so I figured it would be good to add the links in. I therefore updated the API to ensure a call for a bibliographical record also returned a list of all entries that cite the record, and updated the bibliography page to add in a count of these entries and links through to them. I think this works really well – for example if you search for Irvine Welsh you can now find a list of all SND entries that feature a quote from one of his books. The only downside is that some bibliographical records (especially for DOST) appear in thousands of dictionary entries, meaning thousands of links appear on the bibliography page. I’ll need to think about how to handle this, as currently all links to entries just appear as in-line buttons. We actually store the quoted text in the database so we could also display the specific quotes in the bibliography page if we want. I also still need to update things so that the links through to the entries lead directly to the first matching quote.
My demonstration of the new interface on Wednesday went very well. Everyone was very pleased with how the new site looks and functions and were also very pleased with the new bibliography page. We had a good discussion session about further possible updates, and it’s now over to the DSL team to use the new interface over the next few weeks (or possibly months) and then send on any required updates to me. I still can’t really share any screenshots of the new site at this stage, and it’s likely to be several months before it goes live, assuming the DSL team want to go live with it.
I’ve been suffering from what is most likely sciatica over the past couple of weeks, which has made it quite difficult to work. I’ve generally been unable to leave the house much, although the pain eases off as the day goes on and I’m able to move more freely in the afternoons. I’ve been unable to come into the University during this time, which has meant I’ve had to join some meetings online instead of in-person, and I also had to miss the Arts and Humanities developers coffee and catch-up that we’d scheduled for Tuesday this week. I was able to make it across to Edinburgh on Wednesday as I didn’t have to leave the house until lunchtime and the pain wasn’t as bad by then. Unfortunately the journey through exacerbated things and I was in quite a lot of pain the next day. It was impossible for me to sit at my desk for more than a few minutes at a time and I’m afraid I had to take the day off sick.
I was still in a lot of pain on Friday but I managed to work again, and I spent the day catching up with emails and investigating the Hansard data that I last worked with some seven years ago, as Marc and Fraser have a student who wants to work with it. I also spent a bit of time fixing an old Robert Burns resource that had stopped working. The Jame Currie site (https://jamescurrie.gla.ac.uk/) is not one of mine, but had apparently not worked properly since a server update. I fixed a number of issues with the code and got it working again. I’ll need to see about fully overhauling the site at some point as it contains some very useful research data but has a rather ancient interface. I also found some time to return to the migration of the Place-names of Fife data to the new place-names system, working through the historical form sources. There’s still a lot of work that needs to be done on the data before it can be fully migrated, but I’m slowly making progress.






