Week Beginning 10th August 2026
I was back in Glasgow and back to a five-day working week this week, as the summer holiday period drew to a close. I spent a fair amount of time this week continuing to work on the DOST Auld Laws project. Luca has been working on the XML queries I’d specified for the advanced search and I was able to test them out and give feedback on them. These are mostly all working as I’d hoped, which is really great, and I was able to begin working on the front-end aspects of the advanced search, with the aim of connecting all of this into the queries Luca had created.
However, I also had to spend quite a bit of time working with the source files, as the project PI Joanna Kopaczyk-McPherson had realised that two of the eight documents should really be split up into smaller sections. This was no straightforward task, as not only did it require the XML documents to be split into smaller sections, but many other aspects needed to be updated. The document IDs needed to be changed, which mean IDs used for image filenames throughout the documents also needed to be changed, with the filenames of the actual images also then needing to be updated too. For page navigation in the site data is stored in a database and this also needed to be updated, and the XML files stored in eXist for search purposes also needed to be updated. There were twelve steps I needed to follow for each required split, which took some time, but thankfully the process went pretty smoothly and we ended up with 13 documents in the site instead of the original 8.
There were further issues to come, as Joanna has now noticed that the documents need further edits and tweaks, and not just minor changes to text but structural issues such as the insertion of omitted lines. This is going to be very difficult to do as lines are linked to coordinates in the images, which was all handled via the Transkribus tool. We exported the XML files from Transkribus month ago and many major changes have been made to the files since the export so it’s not going to be possible to re-import them. Manually creating new lines in the XML files as they are now will not have the connections through to coordinates in the corresponding image files, so we’ll end up with inconsistent data. It’s not a great situation to be in, and ideally all editing of the documents should have been completed in Transkribus before we exported the files, something I’d mentioned back when I undertook the export process months ago. We haven’t reached a decision on how best to handle this yet and we’ll continue to discuss the options next week.
Despite all of this I did manage to work on the advanced search, creating the advanced search form which, as specified, features a textbox where you can enter some text, a list of documents from which you have the option of selecting the ones you’re interested in and a list of tags that you can either include, exclude or limit your search to, as you can see in the following screenshot:
As of yet I have not implemented the tag limits, but the limit by documents is operational. This connects through to Luca’s new eXist-db queries to perform a search limited by documents and as with the quick search, you can also sort the results by words to the left and right of the term, although I’ll need to get Luca to look into how the KWIC is generated for the advanced search results as they don’t seem to cross line boundaries, unlike the quick search. The advanced search results display information about the documents you’ve selected and if you press on a search result to load the corresponding page the link back to the search results takes you to the right place. There’s still lots to do. The limit by tags is the biggest thing and will probably take quite some time. I also still need to add in an option to cite a specific search result page and add in an option to refine your search, which will remember the options you previously filled in when you return to the search form. An option to clear the search also needs to be added. I’ll continue with this next week.
Also this week I participated in two meetings about the place-names AHRC proposal I’m involved with, and this is coming together very well. We now have an outline proposal completed and pretty much ready for submission.
I also spent a bit of time continuing with the travel routes for the HiMuJe Malabar project, adding in a few more travel routes that had been prepared and made a few tweaks to the XSLT that generates the entry HTML for the new DSL interface, fixing some issues with the layout of the new ‘combs’ sections.
Week Beginning 3rd August 2026
I worked a total of four days over the past two weeks, and was on holiday for the remainder. During this time I had a meeting to further discuss a place-names related AHRC proposal I’m involved with. I can’t really say much more about it at this stage, but the proposal is coming together. I also had to spend some time working with IT Support and Luca to figure out why our local server kept going offline repeatedly. It looks like this was caused by the server getting swamped by requests from one particular source (almost certainly bot or AI) and thankfully IT Support were able to block this, after which the server was stable again. It’s something we’re going to have to keep looking out for in future. I also spent a bit of time working with Luca to get automatic WordPress updates working on our local server, as the way sites had been set up meant that the setting was not working. Luca managed to find a solution to this, which is really great.
I also spent a bit more time working on the Burns Supper Map, creating a record for it on this site (see https://digital-humanities.glasgow.ac.uk/project/?id=156), adding more suppers that had been submitted via the survey and making some requested edits to existing suppers. I also set up access to Google Analytics for the other two members of the project team.
In addition, I investigated an issue with the Scots Syntax Atlas after someone suggested that the linguists’ atlas was looking somewhat blurry. I managed to figure out why this might be the case, although I’m not entirely sure whether this is a new issue or if the markers always looked like that. I contacted the project PI and suggested a couple of updates, but I haven’t heard back yet so will need to wait and see what she says.
I spent most of the remainder of my time working on updates to our new test interface for Dictionaries of the Scots Language and working through the list of outstanding items for the DOST Auld Laws project. For the DSL I completed the updates to the bibliography page that I began working on a couple of weeks ago. I implemented pagination of the entries associated with bibliographical items, with navigation bars appearing above and below the entries, with 20 appearing per page and ‘jump to page’ buttons also appearing, just like with the search results. This works pretty well, but is somewhat cumbersome for someone like Sir Walter Scott, who is referenced in 2861 entries, split over 144 pages. We don’t have this issue with the search results are these are capped at 500 (25 pages) so we might need to think of other ways of handling this.
I also ensured that headword searches that don’t yield any results automatically perform a fulltext search for the term supplied. This works on the live site, but only when both dictionaries are selected. With the new site there are separate quick searches for SND and DOST so the additional search wasn’t being triggered. It is now, as is the advanced headword search for both dictionaries.
I also made tweaks to the DSL’s new regional map based on feedback I’d received – adding in some content where we previously had placeholder text and ensuring the ‘About’ popup didn’t disappear off the bottom of smaller screens and a few other small updates. I then began to look at the ancillary pages and how we can make them look a bit nicer. I spent a bit of time on the ‘Word of the Week’ page and liaised with William Ashford, who is responsible for such content about this, and further updates that were going to make to the ancillary content closer to the launch date of the new site (which will hopefully be in November).
For the DOST Auld Laws project I added the top navigation bar that will link the site in with the SCOTS Corpus and CMSW. I added in copyright information and added facilities to download page images and the XML files for each document. These being up a pop-up asking for users to abide by the license before leading to the actual content, which hopefully won’t be too annoying. I also added in the ‘cite’ popup to all document pages, which took a little time to implement, and added in Google Analytics. I also made the image thumbnails on the document overview pages smaller and placed them in a collapsible section that is closed by default, plus I removed the introduction to the documents page, as this will be covered by the homepage.
I also removed some pages that didn’t have content (e.g. blank pages) from the beginning and end of some of the documents and I added a feature to turn off and on the line highlighting feature. The highlighting feature allows the user to click on a line of text in the image or text for a page and for that line to be highlighted in both the image and the text, which is pretty nice. Unfortunately the line highlighting gets in the way of the image viewer’s zoom and pan functionality on touchscreens, making it a somewhat unreliable and frustrating experience. This new feature removes the option to ‘click’ on a line, meaning pointer events are not intercepted and make their way reliably through to the image viewer, which works much more smoothly.
Also this week I had an email conversation about user feedback and walkthough videos for the STAR resources and booked my accommodation for the DHC conference in Sheffield. Next week I’m back in Glasgow and back working a full week, with summer holidays all over.
Week Beginning 20th July 2026
I worked on Thursday and Friday this week, having taken the other days off as a holiday. Whilst I was away there was an issue with the server on which we host a lot of our important websites, meaning they were unavailable from Sunday afternoon until about 5pm on Tuesday, and I had to spend some time liaising with Luca about what should be done and responding to users and project members who were unable to access the sites. I also had to spend some time once I was back updating everyone on the situation, and on Friday afternoon Luca and I met with Mike Irwin from IT Services to discuss the situation and what we could learn from it. Hopefully we have a plan but we’ll just need to see how things work out in future.
Over the weekend the Burns Supper Map (https://burns-supper-map.gla.ac.uk) had its official launch (thankfully this website is not hosted on the server that encountered issues) at the Burns Birthplace Museum. As I was away on holiday I wasn’t able to attend, but I was kept updated by the project team and it all went very well. After the launch a few more suppers came in, and some existing suppers needed their details updating, for example because their title was not quite right, or their position on the map needed tweaking, or further photos were submitted. I spent some time on Thursday making these updates. I also exported the data for the Old English Thesaurus as CSV files, which are going to be submitted to the Oxford Text Archive, responded to some queries from the Dictionaries of the Scots Language team and made a minor tweak to the parts of speech section of entries on our test server.
Other than my meeting with Luca and Mike during the afternoon, I spent most of Friday working on the travel routes for the HiMuJe Malabar project. I had been sent data for three travel routes to add to the map to test the travel route system I’d previously developed.
I updated my map code to link directly to the data source for the map that is generated from the digital edition and hosted at the University of Jena, so that new updates to places will automatically get pulled in. I also updated the code so that when you select a route from the menu the map displays the full extent of the route, which I think will be very helpful.
Some of the places that appear in the travel itineraries are not yet found in the places file, or are found but do not yet have location data so for now I’ve had to remove these from the travel routes. This isn’t a huge issue though, as for now the routes are really only for test purposes and I can add the missing locations in once they are available.
By the end of the week I’d added the three new routes to the map, although there is still a lot of work that needs to be done. For example, the colours used for the markers and routes are still not finalised and we’ll definitely need to ensure the route colour is not also used for a polygon. Currently Malabar is red, and so is the route, which makes things very confusing. I also need to work on the menu to break it up into sections by type and assign different colours to the routes (or possibly the types). But here is a screenshot showing the full extent of one of the trade routes:
I also had an email conversation with Tom Bartlett about a podcast site he created a while back that he would like to migrate to the University system. I’m going to have a meeting with him about this, hopefully in the next few weeks. I’m on holiday again next week and some of the following week, so it will be a while until my next update.
Week Beginning 13th July 2026
This was a four-day week for me as I’d taken Friday off (and I will also be off for the first three days of next week). I finally managed to assign some time this week to implementing updates to the new Dictionaries of the Scots Language interface based on feedback that had been sent to me earlier this year. I spent most of Monday and Tuesday working on this.
I can’t share any screenshots of the new interface at this stage, but I updated the search results box in the entry page to remove the tab for the other dictionary when performing a search (quick or advanced) for a specific dictionary. This avoids misleading people as it otherwise the tab displays zero results for the other dictionary when in fact it just means the dictionary wasn’t actually searched. Now when a quick search is performed, the other tab heading is replaced by a link to search the other dictionary. Pressing on this performs whatever search you’ve executed (quick or advanced) on the other dictionary.
I also tweaked the font colour of the inactive tabs. I realised that the white text on grey made it look like the tabs were disabled, when they’re not, they’re just inactive. I therefore made the font darker, which I think works a lot better. I then added in a button that scrolls the page to the search / browse box. This appears above the entry header (and in the ‘sticky’ header that appears as you scroll the page) and only appears on narrower screens (where the infobox appears below rather than beside the entry text). Where the entry is in the search results the text is ‘Scroll to results list’ with a down arrow. Otherwise the text is ‘Scroll to browse list’. The DSL team had requested that results term highlighting should be off by default, so I made this change too.
I then began to rework the bibliography page based on feedback. We’ve decided to go with the version of the bibliography page that displays the quotations from any associated entries in addition to the headwords and links through to the entry pages. This required some reworking of the API so that rather than returning individual citations, it brings back entries with each associated citation as part of this. This means that multiple citations for an entry no longer appear as separate items in the list but are grouped by their entry, much like the quotation search results. I also updated the count above the citations to display both the number of entries the item appears in as well as the number of citations (e.g. ‘Cited 718 times in 597 entries’) and each citation also includes its date as well now. However, the order of citations needs to be the order they appear in the entry and not date order, otherwise the links through from the citation may end up taking you to the wrong one in the entry (as I discovered when I set things to date order).
The display of entries and citations is not exactly identical to the search results. There is no sparkline as unfortunately this is not included as part of the bibliography data and I’d need to rework the database and API in order to include them. Also, the quotations are in a larger font than the ones in the quotations search results as they seemed a bit small, and the headword isn’t highlighted in the quotations as this is something performed by the Solr search engine and isn’t available as things currently stand for the bibliographies. I still need to add in pagination, which I didn’t have time to work on this week but will hopefully implement soon.
I also participated in a Teams meeting with the DSL this week that involved their interns reporting back about user engagement and observations about the website. It was very interesting to hear their feedback and it will give us lots to think about as we continue to improve the resource.
Other than working for the DSL, I also made some last-minute updates to the Burns Supper Map before its official launch over the weekend. The resource is now available for anyone to use at https://burns-supper-map.gla.ac.uk. I wasn’t able to attend the launch as I was away on holiday but from what I’ve heard it was a great success.
I also spent a bit of time fixing an issue with the Anglo-Norman Dictionary where certain entries that had an apostrophe in their headwords (e.g. j’) were not loading while others were. The reason for the discrepancy was because some headwords had curly apostrophes and other had straight ones. The straight ones were getting encoded (e.g. “j'”) and were then not getting found. Once I managed to figure this out I was able to fix the issue.
On Wednesday I had a lengthy online meeting for the HiMuJe Malabar project to discuss the interactive travel routes. The meeting lasted about two and a half hours but it was worthwhile as we all have a much clearer idea of how to proceed with the routes now. I just need to wait until the team sends me some initial travel routes using the spreadsheet template I sent them and then I’ll be able to continue my work on this.
On Thursday I participated in a Teams call with colleagues from Nottingham and Cardiff Universities to discuss a place-names proposal that I am likely to be involved with. I can’t really say much more about it at the moment, but it’s all sounding very interesting. I spent a few hours after the meeting writing a document containing some initial thoughts about how the technical infrastructure for the project could work.
I was then off on holiday on Friday and I won’t be back at work again until next Thursday.
Week Beginning 6th July 2026
I returned to work on Monday this week after a lovely holiday. Most of my time this week was spent working for the Dictionaries of the Scots Language, processing a new dataset exported from their editing system and integrating it into the website. The reason this process took considerably longer than usual is that the data had a major structural difference: Entry parts of speech had been moved, rationalised and restructured. Previously parts of speech appeared as elements within <meta> but also appeared embedded in the main entry where they were only tagged with HTML italic tags. There were more than 500 different combinations of parts of speech and things like homonym numbers were mixed in with them too.
The DSL editors have spent a huge amount of time working on the parts of speech to separate them out, standardise them and ensure other data such as homonym numbers are stored in a separate but related manner. The new structure also ensures that entry parts of speech are only stored once in the entry XML using a structure that makes sense semantically rather than for display only.
As the new entry XML in the data export now differed markedly from the earlier structure I then had to rewrite my data processing scripts, update the database and Solr cores, and the API and front-end to deal with the new structure. This was a lot of work. My first step was to create a new table to hold the individual POS data, consisting of a part pf speech and associated ‘hom’, ‘syn’ and ‘infl’ values. I then updated my extraction script so that the existing pos field in the entry table now gets its content from the ‘origpos’ attribute (so we continue to have a record of the original part of speech) and then to populate the new POS table to store each individual pos, including the value, hom, syn and infl data (where applicable) for each entry.
With this update in place I could then (after running a few smaller-scale tests) process every entry in the SND and DOST export files, which resulted in 35,655 parts of speech records being generated for SND entries and 49,828 being generated for DOST.
My next step was to update the entry browse order, which is used to decide in which order the entries appear in the browse pane when viewing entries. This previously used the old POS system to decide in which order entries with the same headword appeared (e.g. so that nouns appeared first). I had to update this to use the new POS system, and also ensure that the POS labels, which are used in the browse pane, the search results and the entry header.
With this update in place I then ran my scripts to process the citations and bibliographies. After that I worked on the Solr cores we use for search purposes. These consist of an ‘entry’ core used for headword and fulltext searching, and a ‘quotation’ core used for searching the individual quotations. Both of these needed to be updated to include new POS fields, which will eventually allow me to add parts of speech filters to the search facilities. I created new Solr cores for each, both featuring new POS fields, after which I would update my script that generates the data to populate the cores to ensure the POS data was included.
With the new cores successfully populated with the new data I then needed to update the API to work with the new POS fields (e.g. so that the labels returned for use in the search results and browse pane use the new POS data). After that I then needed to update the XSLT scripts that process the entry XML files for display to ensure that the new POS structure is displayed when viewing an entry.
At this point all of the updates were still running on my laptop, and with everything in place and tested the final step was to migrate everything to our test server. I don’t have direct access to the server so needed the help of someone with this access in order to complete tasks such as creating and populating the new Solr cores, and thankfully Luca agreed to help out. The process went remarkably smoothly and by the end of Thursday the update was complete, with our test server displaying the new POS fields. I also made a number of other minor updates to the display of entries via the XSLT files that had been requested and it’s now over to the DSL editors to test everything out and make sure all is working as it should.
I spent the remainder of the week dealing with emails I’d received whilst I was away. I also had a meeting with Joanna about the DOST Auld Laws project, which is nearing completion now, and had a meeting with Renu regarding the display of polygons in the map for the HiMuJe Malabar project.
Week Beginning 22nd June 2026
This was a two-day week for me, as I headed off on holiday on Wednesday and will be away until Monday the 6th of July. I mostly spent the two days tidying up loose ends before I went away. This included making some updates to some of the data for the Burns Supper Map that Cleo asked me to do. This was mainly fixing typos and updating the categorisation, plus adding, deleting and replacing a selection of images.
I also wrote scripts to process the edited announcements for the Playbills project. Deven had been working on a spreadsheet version of the data that featured columns noting where announcement text had been edited, where the announcement needed to be deleted, reclassified as a ‘special attraction’ or in some cases where a new announcement needed to be added. After writing and running my script a total of 1836 announcements were updated, reclassified or deleted. I then ran my script that generates complete JSON files for every playbill in the system, featuring all data about associated plays, venues, actors, roles and such things. This took quite a while to execute as a lot of data needs to be processed, but after half an hour or so I had a collection of 1902 JSON files that I sent to Deven so she can run some offline queries for a paper she’s writing.
I spent the rest of my two days continuing to work on the redevelopment of the Mapping Metaphor resource, which will enable both English and Old English metaphors to be viewed in combination. I managed to complete the combined card view, as the following screenshot demonstrates:
The cards feature both main and OE metaphor IDs, counts of lexemes in both E and OE and examples of metaphor from both E and OE. Where a metaphor exists in both E and OE and the E date is later than OE the later date is highlighted in the card in green. I still need to work on the timeline view, the CSV download and the search facilities, plus I need to add in an option to switch between the combined and E / OE specific maps, and also update the user interface, so there is still a lot to do. But not until after my holiday.
Week Beginning 15th June 2026
After being off sick on Thursday and Friday last week I returned to work for a full five days this week, although I was still in the throes of ‘post viral fatigue’ and it wasn’t until Friday that I stopped feeling exhausted. Despite this I still managed to get a lot done for several projects this week. I had been planning on continuing with the redevelopment of Mapping Metaphor towards the end of last week, and as I was unable to do so I decided to focus on this on Monday.
As I looked through the partially updated site I spotted an issue when combining E and OE data: Sometimes the metaphorical connection between two categories has the categories the other way round in the other dataset (i.e. in E we have cat1 and cat2 whereas in OE we have cat2 and cat1). This was resulting in metaphors in the combined view sometimes being treated as different when they’re actually the same, which was affecting some counts and also sometimes making the figure for number of OE lexemes appear next to the wrong category in the card view. I managed to sort this, and I also ensured that if the category order needs to be swapped then the direction is also swapped.
I then spent most of the day implementing the combined table view, and it should be fully operational now. It was extremely tricky to implement this as the script that generates the data for the table view can handle many different data requests and (unlike the visualisation script) also processes individual metaphor connections for display, including directionality and examples of metaphor. There was a lot to update and check, and it was probably not the best of choices of things to work on during my first day back at work. However, I got there in the end and was able to share the update with Wendy and Carole by the end of the day. I still need to update the CSV download option, but that will need to wait for another day.
I also spent about half a day or so on the Playbills project this week. I fixed the few remaining plays that didn’t have links to canonical records so that every play now includes such a link. I then moved onto looking at performers. I wrote a script that outputs a spreadsheet that lists all performers arranged alphabetically by surname with the corresponding play, date of performance and venue. I’m hoping we’ll be able to do something with this to figure out which performers are actually the same person, but with almost 52,000 rows we’ll probably want to try some automated processes rather than figure it all out manually. I’m just not sure how best to proceed with this.
For example, there are 17 ‘Mrs Ashton’ performers, with performances from 1827 to 1834. 12 of these are at the Theatre Royal, Birmingham in 1827, then there’s nothing until 1832 at Bristol (4 performances) then one final performance in 1834 in Edinburgh. Would we consider these all to be the same performer? Or should we treat the Bristol and Edinburgh ones as different people? This will need some discussion with Deven.
I also began looking into generating canonical roles for plays. I wrote a script that for a given canonical play ID returns all plays that have the ID and for each lists its name and all associated roles. The individual play names are links through to the play page, from which it’s possible to access the playbill page and check the original image. At the very bottom of the output is a list of all unique roles with a count of the number of times they appear. This is only an initial version but it has highlighted many issues due to variant spellings (either in the original playbill or due to the AI text extraction). I can further tweak this, for example making all lower case, removing punctuation, removing text in brackets, replacing “M’” with ‘Mac’ or ‘Mc’ to see if these match. We could also use Levenshtein distance to try and find forms that differ by one character, which would match ‘Baillie Nicol Jarvie’ and ‘Balie Nicol Jarvie’ with ‘Balie Nicol Jarvie’ and ‘Captain Thorton’ with ‘Captain Thornton’, for example. We’d have to watch out for false positives, though (e.g. ‘Jane’, ‘Jean’ and ‘Janet’), and also make sure the form that has the most hits is taken as the canonical form, rather than it just being the first form found. But it’s also possible that asking AI to do this might be a better approach.
As I was working on this I noticed the interestingly named role ‘MacSycophant’ in the list, which is definitely an AI hallucination! Looking at the original playbill image it should be ‘MacStewart’. I also spotted that many of the actors are incorrectly assigned, with ‘Mr Harrold’ and ‘Mr Felton’ being swapped and Rob Roy listed as being performed by Miss A Murray instead of Mr Pritchard). Checking the original YAML shows that the issue was with the AI extraction and thankfully it was not any subsequent processing scripts that have introduced the errors.
Following on from this, Deven has suggested that she and some other researchers may manually proofread all of the actors and roles across the entire dataset. This would really help to ensure that the data is more accurate, but it is also a huge amount of work, as there are 1902 playbills containing tens of thousands of performers and roles.
I spent some time proposing a means of editing the data so that the correct updates are made in the database. We would need to generate lists of performers from the database so we don’t lose the work we’ve already done to split the names and assign genders, and presumably the task won’t just involve swapping roles between performers, but will involve role names being edited too (e.g. to fix things like ‘MacSycophant’). I think it’s likely that some issues with performer names that have been extracted incorrectly too will also be spotted and need fixing too. We’ll need to ensure any updates to performers’ names and role names can be tracked to the relevant record in the database using the corresponding ID fields.
To achieve all this I generated a spreadsheet that includes information about the playbill, play, date and venue, with a link to the page for each playbill for checking the original image. The Role data and Performer data then follow on each row. Using this a proofreader could then make changes to the ‘Role Name’ field where required, but also make changes to the performer fields too if anything needs corrected there, and also add in new rows if required.
I then returned to the DOST Auld Laws project, for which I created an initial version of a user interface for the site. This is still just a work in progress and may change depending on feedback given, but below is a screenshot:
I spent most of the rest of the week working on created a system to manage travel routes for the interactive map for the HiMuJe Malabar project. My first task was to create a test route based on data provided by project Co-I Ines. I added this to a spreadsheet template I’d created, and which we’ll hopefully be able to use for future routes.
With this data in place I then wrote a script that converts the spreadsheet into JSON data for use on the map, so if further travel routes are created (one per spreadsheet) with the same structure I’ll be able to add them to the map.
Once I’d completed my updates to the map interface, when you press on the ‘Travel Routes’ section of the map menu each travel route is now listed (there is currently only one). Each route appears with its title and type (we can maybe split and/or limit the list by type in future) and the description given in the spreadsheet. Pressing on the checkbox or the route name adds the route to the map. This currently highlights the relevant locations with a red border and adds a dotted red line connecting each location in the itinerary. We may want to use different colours when multiple routes are added at the same time to help differentiate them.
As the route can get lost in amongst all of the other data on the map I’ve added an option to show or hide unrelated places. If you deselect the ‘Show unrelated places’ checkbox any locations that are not part of the travel route are removed from the map, making it much easier to see the route. The following screenshot shows this:
The legend is removed from the map when the ‘Travel Routes’ map menu is active. This is because we need the space to display the information about a specific stage in the travel route itinerary, as described below. Also, if the legend was visible users may end up removing a map layer (e.g. ‘Temple’) that contains locations that appear in a selected route, which would cause confusion. When viewing travel routes, all of the categorisation layers are set to on for this reason. If you navigate from the ‘Travel Routes’ menu to another menu any selected travel routes are removed from the map, the legend is reinstated and all locations are added to the map.
If you press the ‘Explore’ button for a route this adds the route to the map (if you’ve not already added it by selecting the checkbox) and the map will reposition to display the first stage in the itinerary. This displays an info box in the top right of the map (in the place the legend would otherwise be). This box displays the route title and the information about the first stage in the itinerary. This can include any information about the stage in the itinerary (including duration, if of relevance). I’ve also included information about the number of stages in the route and the number of the current stage (e.g. ‘Stage 2 of 10’), as I figured this would be useful for people. There is also a ‘close’ icon in the top right that closes the info box and ‘Next’ and ‘Previous’ buttons at the bottom that can be used to traverse the travel route. Pressing on one of these buttons repositions the map to the next or previous stage in the itinerary and loads the information about this stage into the info box. Using these options you can step through the travel itinerary and view all of the information relating to it. The following screenshot shows the map with ‘Stage 2’ of the travel route loaded. Note that you can also still open the popup for any location on the route or manually scroll the map between locations.
This is only a first version of the feature and there’s still a lot to do. Displaying multiple routes at the same time is going to require further work, and I also want to add the selected route options to the page URL to enable citation / sharing / bookmarking of specific route selections and also possibly individual stage selections. We may also want to add information about relevant routes to the location popup (in a new tab) so users can tell at a glance which routes a location appears in. I guess adding links from this to the relevant stage in each route would also be useful.
Also, the current system is only set up to work with a purely linear route (e.g. a->b->c). If we are to include routes that branch off (e.g. a->b then b->c and also b->d then c->e and also d->e) then we’ll need to consider how to handle this as the traversal via ‘Next’ and ‘Previous’ buttons would not work.
Also this week I made a few further tweaks to the Burns Supper Map, including updating the site title to include ‘Worldwide’, fixing some typos in the data and adding in some new videos. I’ll be continuing with this next week.
Week Beginning 8th June 2026
I began feeling unwell on Monday this week, but managed to struggle through until Thursday morning, by which point I just couldn’t sit at my desk any more. I was then off work sick on Thursday and Friday.
Despite coming down with something I still managed to get quite a bit done on Monday to Wednesday this week. On Monday I mainly focussed on the Playbills project. I manually sorted all but three of the remaining plays that didn’t have canonical records and asked Deven to further investigate the remaining three. I then added the canonical play data to the output produced by the API and wrote a script that would output all of the playbill data for every playbill as JSON files. This will be used for offline analysis by Deven.
On Tuesday I continued to work on the map for the HiMuJe Malabar project, mostly tidying up some loose ends with the map. I implemented the URL shortener and set up the ‘share / cite’ options so that now if you open a place record and press on the ‘Share / Cite’ tab this loads content, with the citation text being dynamically generated based on the map options that are currently selected. It’s also now possible to press the ‘Share’ button at the bottom of the map menu to share or cite a specific map view as opposed to a specific record.
I also implemented the ‘Reset map’ option on the ‘Home’ menu and pressing on this resets the map to the default view and categorisation type and I finally sorted out the polygons on the map so they should always now appear behind the markers, even when turning layers on and off in the legend.
The last thing I implemented was the ‘Table view’. If you press on this button at the bottom of the map menu it opens a popup containing all of the data in tabular form. I haven’t included the references in this as this is too much data for such a table, but I have included a count of the number of references for each place. Similarly, I’ve only included the latitude and longitude for each place and not the full GeoJSON shapes for the polygons. You can press on a column heading to order the table by that column (press a second time to reverse the order). I think it will be quite a useful view, for example if you know the name of a place but are not sure exactly where it is located you can order the table by placename and find it. Pressing on a placename in the table closes the popup, centres the map on the location and (after an intentional brief delay so you can appreciate where on the map you’re looking at) the relevant record popup opens. I might see about highlighting the place’s marker to make it clearer which is the relevant one in areas where there are many.
I still need to implement categorisation by certainty, but as of yet we don’t have any certainty data in the JSON file so I’m leaving this for now. I’m intending to start on the travel routes soon. We don’t have any routes available yet, but I’m going to create a test route and create the interface for plotting the connections along the route so we can see how this might work.
On Wednesday I continued to work on the DOST Auld Laws project. I managed to fix the issue with multiple results in one line causing the in-page navigation to stop working. This required quite a bit of reworking as the results were identified by line and instead I needed to add in a new way of identifying that a result may be on the same line but is actually a different result. I did this by counting the number of results per line and using this figure in addition to the line ID to track the results, adding this counter to the URL that is used to reach the document page from the search results, with the document page then taking this figure and using it to ascertain which result is the current one and allowing the results traversal on the document page to function.
I then moved onto looking at results page ordering. Previously the search results were ordered by document name, document page and line, but I wanted to add further sort options via a drop-down list to the right of the ‘you searched for…’ box. In addition to the default ordering, this allows you to order by highlighted term (useful for searches like ‘*other’ where the start of terms may differ) and a concordance-like word position. The latter allows you to sort the results by up to 5 words to the left or right of the term, with the selected position highlighted in cyan. The screenshot below shows the search for ‘*other’ ordered by the second word to the left of the term:
This was a particularly large and complex update to implement as many parts of the code had to be rewritten, but I think it’s worth it. The final update I implemented was to ensure that the ‘return to search results’ button in the search bar of the document page now takes you back to the specific results page and retains the ordering you’ve selected rather than taking you to the first page of results and the default ordering.
Also this week I set up a bare-bones WordPress site for Craig Lamont’s new project, and I’ll meet with him over the summer to get this fully set up. I also tweaked the Burns Supper Map to ensure long subtitles for the charts don’t end up overlapping with the ‘hamburger’ menus for each chart and read through the findings of the user survey for the Dictionaries of the Scots Language that I’d been sent, and which was on the whole very positive.
Week Beginning 1st June 2026
The project I spent the most time working on this week was the DOST Auld Laws project, which I hadn’t worked on for a couple of weeks. The last time I worked on the project I’d managed to get an initial version of the quick search working, but there were some issues with it. A phrase search wasn’t working, wildcards at the start of a search term were not working, the full term was not getting highlighted in the results where an <expan> is present in the term, I hadn’t implemented the pagination of search results and I still needed to update the page view to enable results traversal directly from the page when accessed via the search results.
I was struggling somewhat with phrase searching and full term highlighting and ended up going round in circles with ChatGPT for several hours and getting nowhere. I eventually asked my colleague Luca, who has considerably more experience and knowledge of eXist-db and XML document querying, for some advice. Thankfully he was able to come up with a query script that did exactly what I needed it to do, which was a massive help as I was really struggling. This is definitely an example of where a conversing with actual person is much more effective than AI and I really must give Luca credit for the help he gave.
By the end of the week my updates meant that it was possible to use wildcards at the start of a search term, as you can see from the following screenshot that shows the results for ‘*other’:
With Luca’s help, phrase searches now work, and the following screenshot shows the results of a search for “the landis of”:
Again with Luca’s help, the search term in each result is now fully highlighted even when an <expan> is present in the word, so for example a search for ‘pebillis’ previously found two results but only highlighted ‘pebill’ in the results as the term was recorded as ‘pebill<expan>is</expan>’. But now the full term is highlighted.
I also implemented results pagination, with is currently set to display a maximum of 20 results per page. The following screenshot shows a search for ‘r?cht’ with the pagination in place:
And now when you reach a document page that is in the search results (e.g. by selecting it from the search results) a search results navigation bar appears above the document navigation bar, allowing you to navigate to the next or previous search result or return to the full search results, as the following screenshot demonstrates:
I’ll probably add in a ‘clear search results’ button here too, and I still need to fix an issue with the results navigation when there are multiple results on one line. At the moment there is no way for the code to differentiate these so the ‘next’ and ‘previous’ links get stuck. I also need to look into speed issues with the search too. But some good progress has been made this week and I feel much more confident using eXist-db now.
Also this week I spent a bit more time on the Burns Supper Map, adding in videos and images for the recently imported supper records. I also had an email conversation with the Dictionaries of the Scots Language people about parts of speech and new data imports, which will probably be taking place in the next few weeks. I also attended a ‘coffee and catch-up’ with the other developers in the College, which was really valuable as always.
The rest of my week was divided between the Playbills project and the HiMuJe Malabar project. For Playbills I processed a spreadsheet containing around 400 performers that Deven had manually processed last week. I ended up doing a bit more manual tweaking after investigating the appearance of some of the performers in the playbill images, and when I ran the spreadsheet through my import script we then had 52,178 performers in the system, and of these some 52,048 of these have a gender assigned, which I pretty amazing.
I then moved onto processing the plays based on the spreadsheet Deven had worked on before Easter that notes which performances actually feature the same play, even if it is not referred to in exactly the same way. Using this I created canonical records for plays, I set some marked plays as ‘special attractions’ and I deleted certain plays that had been marked for deletion. Before I did this I updated the database so that all tables include an ‘isactive’ field, and updated the API so that it only includes data where ‘isactive’ is set to ‘Y’. Then when it came to deleting plays I didn’t actually delete them but set them (and all associated data such as performers and roles) to ‘isactive = N’. This means if we realise something has been ‘deleted’ that needs to be reinstated I’ll just need to update the relevant fields back to ‘isactive = Y’ rather than having to find and re-insert all associated data.
For the most part the script was successful and result in 261 plays being deleted and 1098 converted to special attractions. It then created 2039 canonical records and 4134 other plays were then assigned to these. There were a few issues with some plays referencing canonical records for other plays that were set to be deleted (so canonical records weren’t created for them), and I passed these on to Deven so she could look into them.
I then created an API endpoint and front-end pages for listing the canonical plays and the details for a selected canonical play. Below is a screenshot of part of the list of canonical plays, ordered by number of plays:
It’s just an initial version that lists the titles, the record type (either play or ‘special attraction’) and a count of the number of associated plays. You can press on column headings to order the table by the column and if you press on a canonical play name you can access a list of plays that are associated with it, as you can see below:
This lists each associated play’s name, date, playbill, venue, location and genres, and you can click on each linked item to reach the relevant page (e.g. the details for a play or the associated playbill page). This is just a work in progress and we’ll probably want to update it, for example to include a genre filter on the canonical plays list, or including access to lists of associated roles and performers. I also updated the playbill page to add links through from play titles to the relevant play page, and I’ve added a ‘Type’ field for each play that either displays ‘Play’ or ‘Special Attraction’.
For the Malabar project I sorted out the legend for source texts, alphabetising the list and ensuring the pane has a maximum width. I also added in the two other categorisation options that I’d included in my specification document: Number of references and placename languages. Number of references categorises the markers and polygons based on the number of times each place is referenced in the source texts. I set up the categories to match the available data (0, 1-5, 6-10, 11-20, 21+) but these can very easily be altered as the data and number of references grow. I think it will prove quite useful to be able to quickly identify the places that are referenced the most in the texts. Below is a screenshot showing the currently available data categorised by number of references:
The other new categorisation option was ‘Placename languages’ and this categorises the places based on the languages of the placename variants included in each record. This will likely need some further work, such as adding in full language names rather than the codes, and possibly filtering out ‘en’ as I’m guessing these would not have been found in the original sources, but it’s still interesting to use the categorisation – for example finding all placenames that have a Hebrew form. Here is a screenshot showing placename languages:
The data itself is still being compiled and I still need to work on the marker colours (and icons) and also to ensure the polygons always appear behind the markers and don’t make the markers unclickable, as is sometimes the case at the moment. Also this week I added the selection of base map, menu section and categorisation type to the URL, meaning it’s now possible to bookmark / share / cite specific views of the map. You can also link directly to a specific record as well.
Week Beginning 25th April 2026
Monday was a holiday this week and on Tuesday I worked on the Burn Supper map. I updated the narrative summary so that it features percentages based on the number of suppers that supplied the data type rather than the overall total. This means that (for example) Whisky now has the percentage 88% rather than 57%. The new summary layout is shown below:
I then ran an import of new suppers, taking the total of suppers up to 1044, and as the facts and figures are all dynamically generated these updated to reflect the new data. I had to do some manual tweaking of the data, for example some countries had typos and needed fixing, plus I needed to process and upload all new images that were associated with the new suppers. Other than future imports of new data I think the Burns Supper map is now pretty much complete.
On Wednesday I focussed on the Playbills project. Last week I’d executed my script to process performers, and this left a few hundred that needed to be manually checked. I started to work on these myself initially, but it was pretty slow going, especially when checking data by querying the database directly. Instead I decided to create a new spreadsheet that includes all of the fields needed for checking (e.g. the role information as well as the performer details) and it should be quicker to edit this. I sent it, along with detailed instructions on how to update the data, to Deven and hopefully she’ll be able to work through the list in an hour or two.
I met with Deven to discuss our next steps for the project during the afternoon, and before this meeting I prepared a list of discussion points. As we would ideally like to know the genders of all roles in the data (so it will be possible, for example, to ascertain when female performers take on male roles) but as roles are often just names without titles assigning gender would require manual checking. Instead I wondered whether we could get AI to assign gender and then leave us with only the trickier ones that are less clear.
As a quick experiment I passed a short list of roles from one play to ChatGPT and asked it to ascertain the gender of each. The prompt I gave was “for the following list of names work out whether each is male, female or unknown: Edmund (the Blind Boy), Stanislaus, Oberto, Rodolph, Kalig, Molino, Starrow, Elvina, Lida” and the response was: “Edmund (the Blind Boy) — Male, Stanislaus — Male, Oberto — Male, Rodolph — Male, Kalig — Unknown, Molino — Unknown, Starrow — Unknown, Elvina — Female, Lida — Female”
So two thirds of roles were correctly assigned a gender and there were no mistakes, leaving one third that would need manual checking, which I think is looking fairly promising.
On Thursday I continued to develop the interactive map for the HiMuJe Malabar project. I sorted out the type / subtype categorisation in the legend to include both types and subtypes, as shown in the following screenshot:
Subtypes are indented within the type now and any places with a type but no subtype (e.g. Sri Lanka) are now appearing. Currently all types and subtypes appear in the legend, even if they have no associated places, mainly so we can see what the full list will look like. It is rather long and I may need to add in a scrollbar, although we may also want to rework the categorisation too – e.g. we have ‘Town’ as both a type and a subtype of ‘Settlement’, plus we have two occurrences of both ‘Hinterland’ and ‘Backwater’.
I was intending the counts beside the types to be a total of all places categorised by the respective subtypes, but some places only have a type and no subtype so the counts represent these instead (e.g. the ‘4’ beside ‘Region’ shows the number of places that have ‘Region’ and no subtype). There are also 5 places that have a subtype within ‘Region’ in addition to this but I’m not sure how best to represent this without confusing people. We could have something like ‘Region (4+5)’ or ‘Region (9)’ but both of these seem a bit unclear to me. I was also thinking of having the checkbox beside each main type select / deselect all subtypes, but if we did this it wouldn’t be possible to just display the main type without its subtypes. These issues need further consideration.
I also implemented the record pop-up that now appears when you press on a marker or polygon, as you can see in the following screenshot:
The popup header displays the ‘preferred name’ for the place and the ‘general information’ tab features the ID, all names and their languages, the category and subcategory (I guess I should standardise this to ‘type’ and ‘subtype’ to avoid confusion) and any supplied description. The ‘references’ tab shows a count of the number of references to the placename in the source texts and the content of the tab lists the filenames and snippets for each reference, with the actual text highlighted in yellow, as shown in the following screenshot:
We should probably have actual titles for the source texts rather than filenames, but these are not included in the output and is maybe something to add in, along with references to specific lines / pages. Another possible issue is that I am aware some of the text will be read right to left and at the moment all text is just displayed as ‘prefix+extract+suffix’. We might need a flag in the data for when the text should instead be ‘suffix+extract+prefix’ (or I guess set the direction to ‘rtl’ in the stylesheet). When we have any images I’ll add these as a further tab, but we don’t have any yet. Also, I haven’t implemented the ‘Share / Cite’ tab yet.
The final thing I’ve implemented is an alternative categorisation for the map, based on the source texts the place is referenced in. In the data I was working with there are only four places that have references, and I’ve added a further ‘No source’ category that all other places are added to, as the following screenshot demonstrates:
I renamed the ‘Place’ menu in the left-hand menu to ‘Place Categorisation’ and there is now an option to switch the categorisation from type to source text. I’ll add in ‘language’ and ‘frequency of reference’ next week, all being well. After I sent an update to the team, Christian, the technical person for the digital edition, informed me that a new output of the data was available that featured many more references and other updates. I therefore replaced the data in the map with the new version and it made a huge difference – the list of cats and subcats is now shorter, there are many more source texts (although this demonstrated that I have some further work to do with the legend for source text categorisation) and many more places with references. I’ll continue with this next week.
During the week Katie Halsey, the PI of the Books and Borrowing project, contacted me to ask for some help in creating some queries of the data for the monograph that she and Matt are writing. These involved working out which books and authors were found at ten or more libraries, and which books and authors were borrowed in every decade from the 1750s to the 1830s. I looked into this on Friday, and it took most of the day to work on it, partially because it’s been a while since I worked with the data and it took some time to remember how everything fitted together. However, I managed to produce the data required data and I sent it to Katie and Matt in four spreadsheets.
Also on Friday I did some work for the Dictionaries of the Scots Language, setting up a new user account that will be used by some interns that are starting with the project over the summer, sorting out access to the Google Search Console and removing the user survey.


















