Week Beginning 16th February 2026

On Monday this week I had a meeting with Joanna Kopaczyk and Pia Lehecka to discuss the DOST source materials transcription project.  The materials have all been transcribed using Transkribus and the project needs a front-end created through which the images and text can be searched and browsed.  A mockup interface has already been very kindly produced for us by Dario Kampkaspar, who created a pathway for converting Transkribus XML to TEI XML (see ‘Page2tei’ here: https://help.transkribus.org/downloading) so hopefully it won’t be too difficult to create something similar for the entire dataset, as task I’ll need to have completed by the end of May.

On Tuesday I had a meeting with Deven Parker regarding her Playbills project.  We met with a PhD student in Computing Science who has been developing the pathway for sending playbill images to AI and processing the files that are returned.  We needed to make a few tweaks to this pathway, and also to figure out how we can run things ourselves and it was a really useful meeting to participate in.  After the meeting Deven and I had a further meeting to discuss the data extraction and processing I’ve been doing, and subsequent data cleaning and amalgamation tasks that will need to be undertaken.  Deven is going to try and proofread and correct the YAML files that were originally generated by the AI as an initial step, and this will mean I’ll need to delete my data and regenerate everything from this updated dataset.  I can’t therefore do much with the data I currently have, but will instead focus on developing the methods for browsing the playbills in the front-end, something I’ll begin working on next week.

I also met with Louis Strange on Tuesday this week to discuss a new proposal he’s putting together.  It was a useful meeting and I gave him some advice about possible ways his data could be presented and used online.  I can’t say much more about the proposal at the moment, but hopefully it will be funded.

Other than meetings, I spent quite a bit of time this week working on the interface for the new interactive map of Burns Suppers.  I set up the basic map interface (using a satellite map with labels) and set up the map menu using the components I created for the Iona place-names project.  For now the side panel has the same colours as the Iona map, but this will be changed in time.

My first task was to ensure that all of the supper records could be loaded into the map, and these work with the clustering tool.  Currently the map marker is the same as the old supper map (red background, knife and fork icon).  I did try using a haggis icon that we found online, but it’s not really going to work as when scaled down to a suitable size it just looks like an indistinct blob.  I also tried using the Font Awesome ‘lemon’ icon as it looked vaguely like a haggis, but I’m not sure this is ideal either.  We’re going to need to give this some further thought.

I added in all of the textual information about the suppers to the record popup and the map is already beginning to shape up quite nicely.  I then added in a facility to allow you to bookmark / share exact views of the map (position and zoom level) using the Leaflet Hash plugin.  Eventually the URLs will also feature filters etc but I haven’t implemented that yet.  But I have implemented the filter options themselves, all of which should now be working.  If you press on the ‘Filters’ menu item you should now see some text about the filters and a drop-down list featuring the various filter types.  I realised my specification document hadn’t included number of guests as a filter type so I’ve added that in too.  One you select a type (e.g. ‘Poems and songs’) the individual filter options are listed, together with counts of the number of suppers that feature them.  You can then press on an option and the map will update to only display the matching suppers.  The screenshot below shows the map with ‘Tam o’ Shanter’ selected.  Note that as of yet the map doesn’t zoom and pan to ensure all matching suppers appear in view – instead the map just stays where you previously had it.  I’m not sure now whether making the map automatically change position is a good idea or not, so it’s something to think about.

I also updated the ‘Home’ menu to add in the title of the resource and to make the buttons work.  The ‘help’ button now opens a popup with some placeholder text while the ‘reset’ button resets the map to include all data and the default position and zoom.  I’ve updated this as previously it was zoomed in on Scotland.  Now the map is zoomed out and shows much of the world (depending on your screen size).  I also ensured the ‘Attribution and copyright’ link in the bottom right now works, displaying a popup with some placeholder text, and I’ve added some text to the ‘Advanced’ menu section, although the advanced filter options are still to do.  I also tested all this on my phone and everything currently in place works fine on it.  Next week I’ll probably work on the advanced filter options, which will allow users to combine different filters.

I spent most of Thursday updating the APIs for various sites, as Luca had spotted some inefficiencies that could be improved.  I’ve now updated all of the APIs that we host locally, but I still need to implement the changes on our externally hosted sites.

On Friday I worked for the Dictionaries of the Scots Language, investigating some non-urgent issues with searching that had been sent to me in November last year.  It turns out that these issues are all sorted with the new DSL interface we’re hoping to launch this year, so I didn’t need to make any changes to anything.  I also made some updates to the layout of some aspects of the new tags on the new DSL website, and the rest of my time was spent making further tweaks to the new DSL website that have been on my ‘to do’ list for a while.

I’d spotted that the search results in the side panel on entries was including blank links that would be highlighted when hovered over and I managed to identify the issue and sort this now.  It turns out I’d already sorted this for searches that involve both SND and DOST, but hadn’t applied the update to individual dictionary searches.  I also created a new version of the bibliography page that lists the actual quotations as well as providing links to the entry pages.  I can’t really share any screenshots of this for now, and I don’t know if we want to include the quotations or not, but I thought I’d create this test version so we can see how it might work.

I also made the links from the bibliographies on both the test version and the main version take you to the actual quotation within the entry, rather than dumping you at the top of the entry.  I used our purpley-pink highlight colour to highlight the citation to make it easy to spot which one brought you to the entry.  This actually took quite some time to implement as a bibliographical item can be cited multiple times in an entry and in such cases the system needs to know which one to link to, but I reckon linking through to the actual quotation from the bibliograph page will be a hugely useful addition.

 

Week Beginning 9th February 2026

I mostly divided my time between three projects this week: Playbills, Burns Supper Map and the Place-names of Armagh.  Unfortunately I was still suffering from the monstrous cold I started with last week and struggled through some of the week, but I still managed to get quite a lot done.

For the Playbills project I wrote, tested and implemented a script that extracts the data from the JSON versions of the playbill files I generated, splits this up and inserts everything into a relational database. The reason I’m doing this is to make it easier to generate canonical records for venues, plays, performers and roles, as it will be much easier to query the data and track records in a relational database.

The data I’ve extracted consists of 1902 playbill records that feature 6434 plays.  These are categorised by one or more of 188 distinct genres (with ‘melodrama’ associated with 2026 plays and ‘melo-drama’ a further 10).  I’ve extracted 185 distinct venues, 49498 performers, 49495 roles and 6083 contributors.

As of yet I haven’t done anything to generate canonical records, which will be the next major step, and I need to discuss things with project PI Deven before I proceed with this.  For example, the role ‘Macbeth’ appears 27 times, with a further three appearances in other strings (not including ‘Lady Macbeth’) e.g. ‘Macbeth’s Last appearance’.  These would need to link to one single canonical ‘Macbeth’ role.  Similarly, there are 28 plays that have ‘Macbeth’ somewhere in their title, with variants such as ‘MACBETH, KING OF SCOTLAND’, ‘Macbeth; King of Scotland’, ‘MACBETH, KING OF SCOTLAND.’ In addition to ‘MACBETH’ and ‘Macbeth’ and these would need to link to a single canonical ‘Macbeth’ play.

There’s also some data cleaning that we should perform, e.g. amalgamating data that doesn’t have the same form but should be the same thing.  For example, there are a lot of possible duplicates in the ‘Genre’ data.  There’s ‘acrobatic’, ‘acrobatic display’, ‘acrobatic performance’ and ‘acrobatics’ all as different genres when presumably these should be the same.

I also still need to work on the performer names to split them into titles, forenames and surnames, and to ascertain gender based on titles.  Venues also need some work as there are many that are the same but have slightly different text, e.g. ‘Royal Theatre, Aberdeen’, ‘Theatre Royal, Aberdeen’ and ‘Theatre Royal Aberdeen’.  There’s the same issue with printers too, although perhaps this isn’t so important.  E.g. ‘Keenes, Kingsmead-Street, Bath’, ‘Keenes, Bath, Kingsmead-Street’ and ‘Keenes, Bath’.  It’s possible that we might be able to get some sort of AI processes to help with such tasks.

We’re also going to have to give some thought about how to handle updates to the data.  I’m generating canonical records, extracting things like performer gender and generating unique identifiers for things like plays in my database, and I’ll be creating new JSON files that incorporate this new data that will then be ingested into Solr for search purposes.  Therefore the data will be quite different to the original YAML files.  When updates need to be made should these then be made to the original YAML files, which would necessitate much regeneration of data, or should the updates be made elsewhere, such as through the database?  I don’t have an answer to this yet, but it’s something we’ll need to consider.

For the Burns Supper Map project I set up the online database for the supper data and have been working on a script that imports the data from the spreadsheets into this database.  I have got everything working for the spreadsheet of the online survey, so my database currently has 308 suppers that include data for 6902 filter options.

What I haven’t been able to do yet is to import the data from the public domain spreadsheet, as this currently contains a lot of inconsistencies in how the data are recorded.  The data in the filter columns (“frequency”, “category”, “toast”, “food”, “style”, “drink”, “entertainment”, “poem”, “music”, “dance”, “dress”) must exactly match the options found in the online form for my import script to work.  This includes capitalisation / case and ensuring that a semi-colon is used to separate multiple items.  I had a meeting with the project RA Cleo on Friday to discuss this, and she’s going to work on tidying things up.

I also wrote a script that posts the address for each record to Google Maps which then returns the latitude and longitude (something we’re going to need in order to pin the records on a map).  The user inputted location data can be somewhat variable, as you might imagine, but Google Maps has generally done a very good job at identifying places from the data, and we can always tweak things once we see the locations on the map.  I’m hoping to start development of the map next week.

For the Place-names of Armagh project I uploaded a large number of place-names that I’d been sent.  We now have 2932 place-names in the system.  I also processed the existing historical forms CSV and this has found historical forms for 1056 of these new place-names.  The new place-names had additional parishes and baronies that were not already in the system and in such cases these have been created, but there are some issues, as the data appears to be somewhat messy at times and will need some cleaning.  For example, there’s a ‘Forkhill’ and a ‘Forkill’ and these may be the same, there are forms with question marks and multiple forms and descriptive text, e.g. ‘Killevy/Partly in Dundonald Parish’ and ‘Armagh?/Eglish?’.  These will all need separated out and fixed as required.

I also spent some time updating the CMS to convert the townland field from a textbox to a list, thus enabling multiple townlands to be associated with a place-name and ensuring each townland is only stored once in the system.  This involved extracting the townlands from all of the 2932 placename records, splitting forms up that have multiple townlands in ‘x or y’ or ‘x / y’ format, storing the unique townlands and then associating the corresponding ones with each placename record.  There are 997 unique townlands (although some of these may need amalgamated) and 3007 connections between townlands and placenames.

I then updated the CMS to replace the existing ‘townland’ textbox with a list of townlands as checkboxes, in the same way as parishes and baronies.  We might need to rethink this, though, as scrolling through 997 townlands to find the right ones takes time.  I also included an option to add a new townland when adding / editing a placename record as I’m guessing there will be more to come.  This should only be used when the townland isn’t already in the list, otherwise we’ll end up with duplicates.

I also made some further updates to the CMS, namely simplifying historical forms so there is just one ‘form’ field rather than separate English and Irish boxes, and adding a flag to record whether the form is a ‘previously suggested form’ or not.  I also renamed the ‘Discovery’ maps to ‘1:50,000’ as this is how the maps tend to be referred to.

Also this week I replied to a couple of emails from the DSL people about future developments and I added a new video to the Seeing Speech resource.  I also investigated an issue with the Books and Borrowing website and discussed the migration of the resource to a new server with the Stirling IT people, and I generated CSV files for all of the survey answers for Speak For Yersel and sent them on to Janine Illian in Statistics, who Jennifer and I met with last week.

Week Beginning 2nd February 2026

I had a bit of a disrupted week this week, as I started feeling unwell on Tuesday morning and ended up off work sick for the rest of Tuesday and Wednesday.  During this time I felt absolutely wiped out and could barely do anything other than sleep, but by Wednesday evening this had developed into a monstrous cold, the likes of which I’ve not had for several years.  Thankfully once the symptoms had moved to my nose and throat my head was a bit clearer and I was able to work on Thursday and Friday, but I was still pretty far from feeling 100%.

I spent most of Monday this week preparing for, travelling to and co-presenting a talk about Speak For Yersel at the Edinburgh Futures Institute with Jennifer Smith.  The talk went pretty well and it was good to meet some of our linguistics colleagues at Edinburgh, plus others involved with the EFI.  I spent some of my other available time reading through and commenting on an AHRC proposal that will involve Glasgow and the Historical Thesaurus that had been sent by Sara Pons-Sanz at Cardiff University, and looking through some further place-name data I’d been sent for the Place-names of Armagh project.

Despite being off work sick on Wednesday I still managed to attend an online meeting for the Burns Supper Map project to discuss the specification document I’d prepared for the project.  This was all very positive and there weren’t any major issues that anyone had spotted whilst reading through it.

For the remainder of the week I spent a bit of time investigating some issues that had been encountered when publishing pure xref entries through the Anglo-Norman Dictionary’s management system.  Certain cross references were not appearing in the published entries despite being in the XML and a bit of investigation uncovered why.  The entries contained cross references to entries that don’t actually exist in the dictionary.  For example, Mars_2 references ‘march’, which is not an entry and respundre_2 references ‘repundre’ which is also not an entry (they both need homonym numbers added).  When xref entries are published the cross references are extracted and stored, and at this point the system checks that the references are valid, and only links to entries that are valid.  It is these that are displayed in the front-end, so even though invalid xrefs may exist in the XML they don’t get displayed.  The ‘preview’ generates its view directly from the XML without checking validity, which is why this view doesn’t match the front-end.  I ran a check and it turns out that there are around 500 xref entries that include a reference to an entry that doesn’t exist, and I passed these onto the editor who will get these sorted.

On Friday I met with Jennifer and Janine Illian, who is the current Head of Statistics, to discuss the Speak For Yersel data and what kind of additional statistical analysis might be possible.  Janine is particularly interested in spatial modelling and has a keen interest in linguistics and it was really great to hear her thoughts about the Speak For Yersel data.  I’m going to send her the data for all survey responses next week so she can experiment with it, and we’ve arranged to meet again later this month.

I spent the rest of my available time this week working on Deven Parker’s Playbills project, working with the YAML files, figuring out how these might be imported into Solr and how we can extract canonical records for things like venues from them.  It turns out that Solr can’t index YAML files (at least not without creating a custom data importer), which is a bit of a surprise.  This isn’t a major issue, though, as I can convert them to JSON, although this also proved to be trickier than I’d anticipated.  Normally I’d use PHP to process data, but PHP also can’t read YAML files, at least not without installing extensions and this process seemed far too convoluted to bother with.  Instead I used Python to convert the files, but this involved a bit of trial and error as I’m not used to Python and it’s bizarre insistence on whitespace being important, and the fact that if you mix up spaces and tabs to create this whitespace the scripts fall over.  I got there in the end, though.

The bigger issue I encountered was with the unit of data that gets indexed.  I’d previously said that we’d index entire playbill files and use ‘playbill’ as the smallest item that gets returned in the search results, but it turns out there are some problems with this, and I think indexing individual plays is going to work better.  I’m still experimenting with the data and Solr’s capabilities, but initial impressions are that it isn’t very good when working with subsets of data within individual files, or more complex queries.  For example, you can search the playbills for the title ‘Macbeth’ and find matching playbills.  But if you combine this with another field that exists in another play in the playbill (e.g. role ‘Jacques Strop’) the playbill record will still be returned.  So even though the role mentioned actually belongs to a different play in the playbill, because both pieces of information exist somewhere in the playbill it gets returned.

With my initial experiments Solr also flattened out the data – all performer names appear in one list per playbill, not separate lists per play, and it’s the same with roles.  Other than the order of the items in the lists, there is nothing to connect the two.  The following screenshot shows one playbill record indexed within Solr (just using Solr’s default post and without customising a schema).  You can maybe see how Solr has flattened things out, resulting in data being lost (e.g. which performer belongs to which play).

I then tried to index the data at play level, adding in a play ID and also any playbill level data (thus ensuring it’s still possible to search for date, venue etc).  You can see the results in the following screenshot, which includes 5 separate records.

Here at least it’s possible to tell which performer / role belongs to which play.  But performers / roles are still only connected by their position in the lists.  Record 5 is a duplicate I made of record 4, but I deleted the ‘role’ text for one performer to see what would happen.  And Solr indexed the record as it was, with 5 performers and 4 roles, so based on list order ‘Miss Newton’ is now ‘Landlord’ and not ‘Marie’, and ‘Mr. Watkins’ now had no role.

After further investigation I realised that it is possible to get sole to properly index nested data (see https://solr.apache.org/guide/solr/latest/indexing-guide/indexing-nested-documents.html) although instructions on how to actually import nested data into Solr are pretty thin on the ground – you can’t just use the default ‘post’ command as this flattens all data.  I ended up following another tutorial (see https://docs.arenadata.io/en/ADH/current/how-to/solr/solr-index-nested-docs.html) and importing the data using the Solr admin interface.  This thankfully worked, as the following screenshot demonstrates.  You can see that individual performers are directly associated with roles.

There’s still a massive amount to do with the data, though.  I need to extract unique venues, plays, performers and roles and assign IDs to them to enable them to be searches for.  I decided that it would be easier to manage such processes via a relational database, so on Friday and mapped out a structure for the playbill data and bean working on an import script that would process the JSON files.  Lots more to do in the coming weeks!

Week Beginning 26th January 2026

Once again this was a week of many different projects.  I spent a fair amount of time on Monday preparing a CV for a Leverhulme bid that Clara Cohen is putting together.  I hadn’t worked on a CV for at least 12 years, so it took some time to look back through everything and prepare the text.  I spent most of the next couple of days on the Burns Supper Map project, with the bulk of this time spent writing a specification document that describes the map I’ll create and the data it will use.  It took quite some time to prepare the document – not just the actual writing of it, but thinking through how the data will be presented and how users will interact with it.  I’d completed a first draft by the end of Tuesday and sent it to the team for feedback.  They are also going to send it on to other interested parties and hopefully they’ll get back to me next week and I can begin work developing the site.

On Tuesday I also met my fellow College of Arts and Humanities developers for one of our coffee and catch-up sessions and it was a good opportunity to hear what they’ve been up to and discuss some of the technical issues we are all currently dealing with.  On Tuesday I also met with Jennifer Smith to prepare for our Speak For Yerself talk in Edinburgh next Monday.  We’re just about there with our preparations and hopefully all will go well.

On Wednesday I set up a bare-bones WordPress site for the Playbills project, as the subdomain and server space I’d requested had come through.  For the moment this does not include Solr (which we’ll need for the searches) and IIIF (which we’ll need for the images), but I’m intending to start developing things locally on my laptop so we don’t actually need these things just yet anyway.  I’ll need some input from the project PI Deven Parker on things like images to use, themes, fonts, logos and colour schemes before we can go live with the initial project website and there’s no real rush to do this.  I’m hoping to start working with the project’s YAML files to extract things like a list of distinct venues next week.

I spent most of the rest of the week working on the Place-names of Armagh project, working with the existing data and creating scripts to import all of the existing data they’d sent me relating to placenames, historical forms and sources into the CMS.  There are now 2999 sources in the system and 233 place-name records, connected to 3412 historical forms.  Almost all historical forms connect through to a source (3408).  There was an issue with a source with ID 203 that was referenced in the historical forms spreadsheet but no source exists with this ID.  It took some time to write and test the import scripts, and it’s possible further tweaking will be required, but I’m pretty happy with how the process went.

I also added in the available grid references.  These were not in the CSV data I’d been sent, but were included in the shapefile data that I was able to load into QGIS.  I was able to export this data as a CSV file from QGIS, and thankfully the IDs in this file corresponded to those of places in the other CSV files I’d been sent, so I was able to join things up and import the grid references.  141 out of 233 placenames have grid references, but what I haven’t had time to do yet is to use this to populate latitude, longitude and altitude.  This is something I’ll need to look into next week.

I also mapped the ‘Status’ codes onto the classification codes in my system.  I’ve added some new classification codes taken from the new data (‘Ro’ for road system, ‘X’ for ex nomine, ‘M’ for minor place, ‘D’ for district).  Other status codes have been mapped onto existing classification codes.  ‘H’ has become ‘R’ (relief), ‘V’ has become ‘S’ (Settlement).  I haven’t imported ‘DY’ as this didn’t seem to fit with the rest of the codes.

I also imported all Parish, Barony and Townland associations for each place.  Some places have a different parish in the 1865 and 1961 columns and in such cases the 1865 parish is associated as a ‘former parish’.  I also imported the map sheets.

On Friday I had a useful meeting with the Armagh team where I talked them through the data in the CMS and we discussed some of the issues that cropped up.  I now have a list of updates that I’ll need to make to the data structures and the CMS, including trying to automatically extract Irish names and translations, renaming ‘Discovery’ maps to ‘1:50,000’, removing the separate language fields from the historical forms, ensuring Townlands can be selected from a list, as with parishes and baronies, and adding in a new ‘Previous suggested form’ Y/N field to the historical forms.

Also on Friday I made a couple of updates to the Anglo-Norman Dictionary to ensure that ‘M.E.’ appears as ‘English’ in the entry and search pages.  I also made a couple of minor updates to the Dynamic Dialects site and helped sort out an access issue that an RA was having with one of the project websites.