Category: Anglo-Norman Dictionary
Week Beginning 25th August 2026
This week I continued with the interactive map for the HiMuJe Malabar project and made some good progress with its development. All subtypes now have an associated icon that is used on the map when the ‘type/subtype’ categorisation is selected, and I also reworked the colours of these, as you can see in the following screenshot:
After last week’s frustrations with the map legend, I reworked it to divide it a little better into sections by type, which you can also see in the above screenshot. When a location popup is opened the corresponding subtype icon also now appears in the popup header. I also added an option to the ‘Home’ menu to make all marker labels permanently visible on the map (until you turn them off again). With all locations on this does clutter up the map, but I think it’s a useful option for people who are looking for a place but aren’t entirely sure where it is. With labels on you can see the names of all places without having to hover over each marker in turn. This works better when the map is filtered in the legend. I also intentionally set the map so that I you set the labels to on and open the travel route menu the labels are removed as the travel routes would be very difficult to see with the labels on.
I also updated the colours for the other categorisation types to hopefully make them work a little better. When the ‘number of references’ categorisation is selected the marker colours are now a gradient. It’s also now possible to cite a specific stage in a travel route. This appears in the citation text, e.g. “HiMuJe Malabar Digital Map: Watercolor map, travel routes menu, showing the travel route ‘Brahmins’, with unrelated places hidden, viewing stage 2 of the route, placenames categorised by language. 2026. In HiMuJe Malabar. Glasgow: University of Glasgow. Retrieved 28 August 2026 and following URL opens the route at Stage 2. This option will mean I’ll be able to update the location popup to add a new tab that lists all of the stages in routes that the place appears in, and give links to each of these. I haven’t done this yet, though.
Also this week I spent quite some time working for the Dictionaries of the Scots Language. This included attending an online meeting to discuss the structure of the new site, at which some fairly major revisions to the new design were proposed. I also engaged in a couple of email conversations about the updates that still need to be undertaken before the hopefully launch the new site later this year. In preparation for all of this work I also went back through all of my emails from the DSL team and my notes to create a bit ‘to do’ list containing everything that needs done. It’s going to be a busy couple of months.
Also this week I met with Craig Lamont and his RA Emily Hay to discuss the Life Writing project. They’re going to be compiling a bibliography of Scottish life writing and I’m going to develop the systems for this. I did look into using existing bibliography software but nothing I looked at did exactly what they are after so I’m going to create a simple CMS for them to use myself, with a front-end consisting of search and browse options and a map interface sometime next year. I spent a bit of time after the meeting going through the sample spreadsheet they had prepared and sending a list of questions about the data.
I also met with Deven Parker this week to discuss an AHRC proposal she’s involved with. I can’t say too much about this, but it’s related to the work we’re doing on the playbills and I would be involved in some technical capacity. I also had a chat with Garrick Allen about a proposal he’s trying to get funding for after an earlier AHRC submission was unsuccessful. I read through the documentation and gave feedback, and I’ve also put Garrick in touch with my colleague Luca Guariento who is probably going to provide the technical assistance.
Other tasks I carried out this week included investigating an issue with the Anglo-Norman Dictionary (which turned out to be an issue with the data and not the system), making some final tweaks to the Burns Supper Map before the project RA finished working on the project, and generating some Books and Borrowing data for a publication Katie Halsey is working on.
Week Beginning 13th July 2026
This was a four-day week for me as I’d taken Friday off (and I will also be off for the first three days of next week). I finally managed to assign some time this week to implementing updates to the new Dictionaries of the Scots Language interface based on feedback that had been sent to me earlier this year. I spent most of Monday and Tuesday working on this.
I can’t share any screenshots of the new interface at this stage, but I updated the search results box in the entry page to remove the tab for the other dictionary when performing a search (quick or advanced) for a specific dictionary. This avoids misleading people as it otherwise the tab displays zero results for the other dictionary when in fact it just means the dictionary wasn’t actually searched. Now when a quick search is performed, the other tab heading is replaced by a link to search the other dictionary. Pressing on this performs whatever search you’ve executed (quick or advanced) on the other dictionary.
I also tweaked the font colour of the inactive tabs. I realised that the white text on grey made it look like the tabs were disabled, when they’re not, they’re just inactive. I therefore made the font darker, which I think works a lot better. I then added in a button that scrolls the page to the search / browse box. This appears above the entry header (and in the ‘sticky’ header that appears as you scroll the page) and only appears on narrower screens (where the infobox appears below rather than beside the entry text). Where the entry is in the search results the text is ‘Scroll to results list’ with a down arrow. Otherwise the text is ‘Scroll to browse list’. The DSL team had requested that results term highlighting should be off by default, so I made this change too.
I then began to rework the bibliography page based on feedback. We’ve decided to go with the version of the bibliography page that displays the quotations from any associated entries in addition to the headwords and links through to the entry pages. This required some reworking of the API so that rather than returning individual citations, it brings back entries with each associated citation as part of this. This means that multiple citations for an entry no longer appear as separate items in the list but are grouped by their entry, much like the quotation search results. I also updated the count above the citations to display both the number of entries the item appears in as well as the number of citations (e.g. ‘Cited 718 times in 597 entries’) and each citation also includes its date as well now. However, the order of citations needs to be the order they appear in the entry and not date order, otherwise the links through from the citation may end up taking you to the wrong one in the entry (as I discovered when I set things to date order).
The display of entries and citations is not exactly identical to the search results. There is no sparkline as unfortunately this is not included as part of the bibliography data and I’d need to rework the database and API in order to include them. Also, the quotations are in a larger font than the ones in the quotations search results as they seemed a bit small, and the headword isn’t highlighted in the quotations as this is something performed by the Solr search engine and isn’t available as things currently stand for the bibliographies. I still need to add in pagination, which I didn’t have time to work on this week but will hopefully implement soon.
I also participated in a Teams meeting with the DSL this week that involved their interns reporting back about user engagement and observations about the website. It was very interesting to hear their feedback and it will give us lots to think about as we continue to improve the resource.
Other than working for the DSL, I also made some last-minute updates to the Burns Supper Map before its official launch over the weekend. The resource is now available for anyone to use at https://burns-supper-map.gla.ac.uk. I wasn’t able to attend the launch as I was away on holiday but from what I’ve heard it was a great success.
I also spent a bit of time fixing an issue with the Anglo-Norman Dictionary where certain entries that had an apostrophe in their headwords (e.g. j’) were not loading while others were. The reason for the discrepancy was because some headwords had curly apostrophes and other had straight ones. The straight ones were getting encoded (e.g. “j'”) and were then not getting found. Once I managed to figure this out I was able to fix the issue.
On Wednesday I had a lengthy online meeting for the HiMuJe Malabar project to discuss the interactive travel routes. The meeting lasted about two and a half hours but it was worthwhile as we all have a much clearer idea of how to proceed with the routes now. I just need to wait until the team sends me some initial travel routes using the spreadsheet template I sent them and then I’ll be able to continue my work on this.
On Thursday I participated in a Teams call with colleagues from Nottingham and Cardiff Universities to discuss a place-names proposal that I am likely to be involved with. I can’t really say much more about it at the moment, but it’s all sounding very interesting. I spent a few hours after the meeting writing a document containing some initial thoughts about how the technical infrastructure for the project could work.
I was then off on holiday on Friday and I won’t be back at work again until next Thursday.
Week Beginning 4th May 2026
This was a four-day week as Monday was a bank holiday. I spent most of Tuesday and Wednesday continuing to work on the DOST Auld Laws project, working with the XML files. The XML files generated by Transkribus and exported by the tool as TEI XML contained many elements that were not valid TEI elements, such as the custom <Aitken> element that had been applied to notes added by A J Aitken. My first task of the week was to write and apply transformations to the XML files to convert them into fully valid TEI. I achieved this using XSLT, which is a language I find very unintuitive and frustrating to work with, no doubt exacerbated by the fact that I don’t work with it very often. I struggled to get any transformations to run initially, and ended up turning to AI to figure out why my scripts were not working. In this instance AI proved to be extremely useful as it identified what the problems were (e.g. I hadn’t declared the correct namespace or used it when writing the rules) and really helped me to understand how everything fitted together. I still wrote the bulk of the code myself, but AI was very helpful in identifying errors or issues. By the end of Tuesday I had generated (and checked) a collection of TEI files that successfully validated in Oxygen, which was a good milestone to reach.
On Wednesday I then worked on the front-end for the project, figuring out how to transform the valid TEI XML into HTML for display in the ‘text and image’ and ‘text only’ views of document pages. As this was once more using XSLT I enlisted the help of AI to figure out specific issues that I was unfamiliar with. The biggest of these was how to pick out and process one single page from a document’s XML file based on the <pb/> element. I had no idea how to achieve this, and despite this being a fairly fundamental issue when processing TEI documents I didn’t manage to find any useful information online. However, AI came up with a solution (and equally importantly an explanation) in a few seconds and I was then able to incorporate this into my code, transforming the contents of one page, whose ID was passed to the XSLT script as a parameter, to HTML with a variety of styles applied to the various elements.
I also updated the ‘click on a line in the image’ feature, and it’s now possible to deselect the line if it’s already selected – previously once you’d highlighted a line you couldn’t get rid of the highlighting, only move it to a different line but now if you press on the highlight it’s removed. The lines are also now connected to the text view – pressing on a line in the image also highlights the corresponding line in the text. You can also press on a line in the text to highlight it and the corresponding line in the image.
I also updated the height of the text pane so that it matches the height of the image pane and if the text is longer the pane scrolls. This ensures that if the text is very long it’s still possible to see the image, rather than having the entire page scrolling, which may result in the image not being visible when you’re at the end of the text. It is how we did things in Books and Borrowing, but having a scrollbar in a section of the page in addition to the browser’s scrollbar can be annoying for some people so I might revert to the previous layout depending on feedback. Here’s a screenshot showing the transformed text, a highlighted line and some of the formatting that’s been added:
Also this week I continued to work on the Place-names of Armagh project. I wrote and executed a script that generated Irish grid references and ITM values for all places based on their latitude and longitude, a task that I completed using the ‘GridRefUtils’ scripts as detailed here: https://www.howtocreate.co.uk/php/gridrefapi.php. I’ve used these scripts before on previous place-name projects and they’ve been hugely useful. I then ran a further script to generated altitude for the place-names by connecting to the Google Maps API, so we now have complete geospatial data for all of the places (other than the 386 that didn’t include Easting and Northing data in the original spreadsheet). I also added in a new parish and barony in a different county that one place-name requires, and had discussions with the team about the splitting of analysis data into Irish forms and translations. Removing the Irish forms from the translation field is going to be rather tricky to automate as the Irish forms often form part of the translation. It’s looking like I’ll need to automate the transformation of some of these, with the rest then requiring manual intervention.
I also continued working on the redevelopment of the Mapping Metaphor resource to create a combined English and Old English map, something I began last week. This week I managed to get the combined view of the drilldown of the visualisation working. The counts represented by the yellow circles also use the combined data, and these are also displayed in the pop-up card, for example, if you select 1K02 Creation and press on the yellow line or circle for 1B Life. In the combined map there are connections to 6 categories in Life whereas there are 5 in the E map and 2 in the OE map (one of the OE ones is also present in the E map, which is why the combined total is 6). The combined pop-up card also now lists the number of OE lexemes in addition to the E lexemes, for example “1K02 Creation 1088 lexemes / 127 OE lexemes”.
I still need to implement the combined visualisation view of the search results, and also the card pop-up between individual categories, which is going to take quite some reworking (possible direction changes, combined example lexemes, timeline updates etc). I’ll hopefully find some time to continue with this next week.
Also this week I processed and added some more images to a few Burns Suppers, had a chat with Garrick Allen about a research project he’s wanting me to be involved with later in the year, and contacted Lindsay Balfour about a proposal she’s writing that will have some technical requirements. I also made a couple of tweaks to the Thesaurus of Old English website after Fraser go in touch with some suggestions, made some tweaks to the survey popup on the Dictionaries of the Scots Language website, gave some mapping advice to Renu of the HiMuJe Malabar project, and created an alternative version of one of the Anglo-Norman Dictionary’s textbase documents that strips out all Latin text.
Week Beginning 23rd March 2026
I continued to mainly work on the Playbills and Burns Supper projects this week, but also spent some time on several other projects. For the Playbills project I re-ran my venue and printer merge scripts as previously the scripts only merged records and didn’t update their names and locations. I hadn’t realised that Deven was going to edit some of these, but thankfully we spotted this and my updated scripts incorporated the edits.
I also began work on a script that will create canonical records for plays. These records will track the appearance of plays across multiple playbills, even if the actual names of the plays differ. Deven has been updating a spreadsheet to note which plays are actually the same and I’ll be using this to create the canonical records. Unfortunately Deven encountered a bit of an issue with Teams, which resulted in her losing a tranche of updates she had made, and this set us back a bit. I have the new database and script in place, so hopefully I’ll be able to run the data through this next week.
I then began to tackle the performers – splitting the names up into constituent parts, reformatting them and assigning gender. It all went pretty well, and my script successfully processed 48,227 performers (34420 male, 13807 female), leaving 1271 that will need further work. The majority of the problematical rows are due to there being multiple performers and it should be possible to refine my script to further process these and create separate performer records for each. This is something I’m going to look into next week. The remainder are either blank, have text such as ‘[Actor not named]’ or are problems with the AI output (e.g. ‘celebrated clowns’ is not a performer name).
For the 48,227 that have been processed my script extracts the first word of the full name and checks it against lists of male and female titles. If the title is in one of the lists the script processes the performer, assigning it the relevant gender and removing bracketed text and saving this in a separate field (any forms of ‘junior’ and ‘senior’ such as ‘Jun.’ are also extracted and stored in this field as standardised ‘Junior’ and ‘Senior’ text).
Some names have exclamation marks and these are removed, and names are converted to lower case with upper case letters for each word. If the name (minus title and other text) is multiple words the final word is treated as a surname and the rest as forenames, otherwise the name is set as surname. Finally, if the surname starts with Mc, M’ or O’ the following letter is capitalised. In the case of M’ this is also standardised to Mc. Note that ‘Mac’ is not changed as names like ‘Macauly’ are not written as ‘MacAuly’ (at least, not usually!).
With the names split like this we’ll be able to have alphabetical lists of performers, which will be useful, as will having gender. It will also help with identifying the same performer across multiple plays.
I also spent a bit of time researching the locations of venues, so we have a few to use for test purposes when it comes to developing a map-based interface to the data. Unfortunately this proved to be quite difficult. Google is pretty hopeless when searching for historical theatre names and just fills its results with random current ‘theatre’ links and ‘book tickets now’ adverts when trying to find actual historical information. It was all very frustrating and I’m sure Google never used to be this rubbish. I did manage to find the locations for around 20 venues, which is a good start, but there are still around 80 left to do.
For the Burns Supper project I mostly spent my time processing data. This included uploading images that had been converted from PDF files and associating them with supper records and running another import of the survey and public domain spreadsheets. This resulted in a further 141 suppers being added to the system, and I also took the opportunity to create a document that lists the steps needed to import new suppers, and this documentation should ensure that future data imports are easier to manage. As part of this process I also dealt with new images for the suppers, running scripts to convert, resize and rename image files and create image records in the database. I also manually split one supper record into three distinct records as it actually covered three different events at different locations.
I also added in links to videos that RA Cleo had uploaded to YouTube. It was a little trickier than I anticipated to get the embedded videos working properly. There were then some issues with the videos not taking up the correct amount of space in the carousel, but I rectified this. I then noticed that the carousel’s ‘next’ and ‘previous’ buttons were obscuring the YouTube player buttons, so if you tried to enter full screen or pressed the play button you’d just change slide. I managed to fix this, but needed to tweak it further when I realised Chrome-based browsers display a different player to Firefox.
There was also an issue with the YouTube player displaying thumbnails of other, entirely unrelated videos in the player when you paused the video. I was shown videos of Donald Trump as a young man and a video about Epstein, which was all pretty awful. Thankfully I found a way to disable this. I then noticed that if you didn’t manually pause a video or watch it to the end before either navigating to another slide in the carousel, selecting one of the other tabs in the record popup or closing the popup, the video (although now hidden) kept playing so you’d still hear the audio. This was clearly no good and I managed to sort it, although my first attempt at a fix didn’t work when a record had multiple videos (it does now). So eventually all was working correctly and we now have 53 videos that are accessible as part of the map.
I also added in a ‘Random supper’ feature to the ‘Home’ tab. This navigates to and (after a short delay) opens the record for a random supper, based on your currently chosen filters, so for example if you have filtered the map to only show suppers that have images or videos you can press the button to load a random supper that has media.
Also this week I had a meeting with Ophira and Renu about the HiMuJe Malabar interactive map. It’s been a while since I met with them as they’ve been focussing on adding place data to the digital edition, which is being overseen by a partner institution. However, we’re reaching the point where the map needs to be developed. We had a productive meeting, discussing the data and how it might be represented on a map. Next week I’m hoping to write a brief, non-technical specification document so we all know what will be developed.
Also this week I ran a few queries and outputted a spreadsheet of data for the Anglo-Norman Dictionary, and I made some updates to the map of dialect areas for the Dictionaries of the Scots Language based on feedback I’d been sent a while back. I think this map is nearing completion now, although I still need to develop region search and filtering for the DSL’s advanced search, which this map will then connect to.
Week Beginning 2nd March 2026
I spent a lot of this week continuing to work on the Playbills project. Last week I began working on a very rough first draft of a front-end through which the playbills data can be browse, and this week I completed this but creating the API call and corresponding page that displays the full data about a specific playbill. As of yet I’ve not included the actual playbill images, so for now there’s just a placeholder section where the image will fit. There is still a huge amount to do for the project and this is just a first draft of the ‘browse and view’ functionality only with practically no time spent on the user interface, but here’s an example of the work in progress:
We still need to do a lot of work with the data, such as amalgamating records for things like genres and venues that have similar but not identical text, plus making canonical records for things like plays. Until this work is completed I can’t really do much more with the front-end, as we need this data before I can generate the Solr search index, and there’s the possibility that all of the existing data will need to be wiped and regenerated as proofreading progresses.
Towards the end of the week I began to download all of the playbill images from OneDrive, a process that took a long time as the downloads kept quitting with an error whenever the ZIP file became larger than around 10Gb – instead I had to download the images in smaller batches. I also realised that a lot of the images are in the HEIC format, which is an Apple-specific format that can’t easily be opened in Windows and is unsuitable for use in websites. At the moment I’m unable to even view the images and it’s going to take some time to figure out firstly how to open them and secondly how to batch convert them to a more widely supported format.
On Tuesday this week I met with Wendy Anderson and Carole Hough to discuss updates to the Mapping Metaphor resource. Currently the resource is divided into separate sections for Old English and the rest of English and we’re hoping to amalgamate these (as well as continuing to give options to focus on one or the other). As well as attending the meeting I spent most of the day preparing a document that discusses the work that will need to be carried out and some of the issues that will need to be addressed. We discussed a lot of these issues at the meeting and the document will prove useful when I come to begin the work, which I’m hoping to do after Easter.
I also spent some time this week working on a few items for the Anglo-Norman Dictionary. A few weeks ago I updated the entry publication workflow so that entries that are cross references are checked to ensure that their references actually refer to entries that exist in the system and the editor Geet had spotted a couple of instances where cross references were continuing to be flagged even though they were main entries rather than cross references. I checked the data and this was happening because both xref and main entries with the slugs are active in the database. I think this may have been caused by the issue of uploaded entries not having the existing ID in the filename, something we discussed a few weeks ago, and I manually deactivated the xref entries, which got rid of the warnings.
I also ran a query to identify duplicate active entries in the system and there were 23 of them. 22 of them were cross references to single entries so it was pretty easy to deactivate the duplicates. One of them had one active version that is a cross reference to four entries and another that only has a single cross reference to a single, further entry, and Geert created a new amalgamated version after I flagged it. The remaining duplicate had one live entry that’s a main and another live entry that’s an xref, but the xref version was malformed so I deleted it.
I also spent a bit of time creating a new cross-reference checker, this time rather than focussing on purely xref entries the new checker instead finds all xref elements in main entries and checks whether they actually link through to an active entry in the system. This script takes a while to run as it needs to check the XML of several tens of thousands of entries, but it’s going to be a useful tool. Currently it has identifies 880 broken xrefs in main entries and Geert can use the output to decide what the do about these. It is something that will more than likely require manual editing rather than any batch processing, unfortunately.
The rest of my available time this week was spent on the Burns Supper map, which is coming along pretty nicely. I’ve updated the records to include information about the source of the data. If it’s a survey response the text is ‘Information about this supper was submitted as a survey response.’. If it’s a PD supper the text is ‘Information about this supper was gathered from publicly available sources.’ I also added a new filter to the advanced filters that allows you to select the source. I figured this would be sufficient, rather than adding another option to the standard filter options. I also updated the map to move the side panel to the right, both to differentiate the site a little from the place-names maps and also because when zoomed out the panel only now obscures parts of Russia (and Japan, unfortunately) rather than North America.
I then moved onto a larger task: getting the map to remember filters and open records so people can bookmark / share / cite specific views of the data. This has been a pretty major update that required a lot of changes to be made throughout the code, but it’s going to be a very useful addition.
Whilst working on this I realised that the advanced filter popup wasn’t working properly in Chrome-based browsers – the popup was unexpectedly closing in Chrome when filter options lower down the form were clicked on. It would appear that there is some kind of conflict with the UI framework I was using for the filter buttons and the map popup, because links that didn’t use this framework worked fine. I’ve therefore replaced the UI framework checkbox buttons with my own and the form now works perfectly in Chrome. The buttons mostly look the same, but the ‘ticks’ are now the browser’s default ticks and can’t be styled. This isn’t a big issue, though.
I also fixed an issue whereby opening records from the table view wasn’t adding the record ID to the URL, and I also implemented the short URLs and the ‘share’ options. If you press on the ‘Share’ button in the side menu this will display a popup containing some buttons linking to sharing platforms, and the citation styles are also displayed and feature a description of the contents of the map based on filters applied. I also updated the record popup to add the ‘details’ to a separate tab labelled ‘Further Details’, with the other textual data in a tab labelled ‘General Information’, and added in a ‘Share this record’ tab, which works in the same way as the map share option. There is still a lot to do for the resource, but I feel like I’m making good progress.
Week Beginning 2nd February 2026
I had a bit of a disrupted week this week, as I started feeling unwell on Tuesday morning and ended up off work sick for the rest of Tuesday and Wednesday. During this time I felt absolutely wiped out and could barely do anything other than sleep, but by Wednesday evening this had developed into a monstrous cold, the likes of which I’ve not had for several years. Thankfully once the symptoms had moved to my nose and throat my head was a bit clearer and I was able to work on Thursday and Friday, but I was still pretty far from feeling 100%.
I spent most of Monday this week preparing for, travelling to and co-presenting a talk about Speak For Yersel at the Edinburgh Futures Institute with Jennifer Smith. The talk went pretty well and it was good to meet some of our linguistics colleagues at Edinburgh, plus others involved with the EFI. I spent some of my other available time reading through and commenting on an AHRC proposal that will involve Glasgow and the Historical Thesaurus that had been sent by Sara Pons-Sanz at Cardiff University, and looking through some further place-name data I’d been sent for the Place-names of Armagh project.
Despite being off work sick on Wednesday I still managed to attend an online meeting for the Burns Supper Map project to discuss the specification document I’d prepared for the project. This was all very positive and there weren’t any major issues that anyone had spotted whilst reading through it.
For the remainder of the week I spent a bit of time investigating some issues that had been encountered when publishing pure xref entries through the Anglo-Norman Dictionary’s management system. Certain cross references were not appearing in the published entries despite being in the XML and a bit of investigation uncovered why. The entries contained cross references to entries that don’t actually exist in the dictionary. For example, Mars_2 references ‘march’, which is not an entry and respundre_2 references ‘repundre’ which is also not an entry (they both need homonym numbers added). When xref entries are published the cross references are extracted and stored, and at this point the system checks that the references are valid, and only links to entries that are valid. It is these that are displayed in the front-end, so even though invalid xrefs may exist in the XML they don’t get displayed. The ‘preview’ generates its view directly from the XML without checking validity, which is why this view doesn’t match the front-end. I ran a check and it turns out that there are around 500 xref entries that include a reference to an entry that doesn’t exist, and I passed these onto the editor who will get these sorted.
On Friday I met with Jennifer and Janine Illian, who is the current Head of Statistics, to discuss the Speak For Yersel data and what kind of additional statistical analysis might be possible. Janine is particularly interested in spatial modelling and has a keen interest in linguistics and it was really great to hear her thoughts about the Speak For Yersel data. I’m going to send her the data for all survey responses next week so she can experiment with it, and we’ve arranged to meet again later this month.
I spent the rest of my available time this week working on Deven Parker’s Playbills project, working with the YAML files, figuring out how these might be imported into Solr and how we can extract canonical records for things like venues from them. It turns out that Solr can’t index YAML files (at least not without creating a custom data importer), which is a bit of a surprise. This isn’t a major issue, though, as I can convert them to JSON, although this also proved to be trickier than I’d anticipated. Normally I’d use PHP to process data, but PHP also can’t read YAML files, at least not without installing extensions and this process seemed far too convoluted to bother with. Instead I used Python to convert the files, but this involved a bit of trial and error as I’m not used to Python and it’s bizarre insistence on whitespace being important, and the fact that if you mix up spaces and tabs to create this whitespace the scripts fall over. I got there in the end, though.
The bigger issue I encountered was with the unit of data that gets indexed. I’d previously said that we’d index entire playbill files and use ‘playbill’ as the smallest item that gets returned in the search results, but it turns out there are some problems with this, and I think indexing individual plays is going to work better. I’m still experimenting with the data and Solr’s capabilities, but initial impressions are that it isn’t very good when working with subsets of data within individual files, or more complex queries. For example, you can search the playbills for the title ‘Macbeth’ and find matching playbills. But if you combine this with another field that exists in another play in the playbill (e.g. role ‘Jacques Strop’) the playbill record will still be returned. So even though the role mentioned actually belongs to a different play in the playbill, because both pieces of information exist somewhere in the playbill it gets returned.
With my initial experiments Solr also flattened out the data – all performer names appear in one list per playbill, not separate lists per play, and it’s the same with roles. Other than the order of the items in the lists, there is nothing to connect the two. The following screenshot shows one playbill record indexed within Solr (just using Solr’s default post and without customising a schema). You can maybe see how Solr has flattened things out, resulting in data being lost (e.g. which performer belongs to which play).
I then tried to index the data at play level, adding in a play ID and also any playbill level data (thus ensuring it’s still possible to search for date, venue etc). You can see the results in the following screenshot, which includes 5 separate records.
Here at least it’s possible to tell which performer / role belongs to which play. But performers / roles are still only connected by their position in the lists. Record 5 is a duplicate I made of record 4, but I deleted the ‘role’ text for one performer to see what would happen. And Solr indexed the record as it was, with 5 performers and 4 roles, so based on list order ‘Miss Newton’ is now ‘Landlord’ and not ‘Marie’, and ‘Mr. Watkins’ now had no role.
After further investigation I realised that it is possible to get sole to properly index nested data (see https://solr.apache.org/guide/solr/latest/indexing-guide/indexing-nested-documents.html) although instructions on how to actually import nested data into Solr are pretty thin on the ground – you can’t just use the default ‘post’ command as this flattens all data. I ended up following another tutorial (see https://docs.arenadata.io/en/ADH/current/how-to/solr/solr-index-nested-docs.html) and importing the data using the Solr admin interface. This thankfully worked, as the following screenshot demonstrates. You can see that individual performers are directly associated with roles.
There’s still a massive amount to do with the data, though. I need to extract unique venues, plays, performers and roles and assign IDs to them to enable them to be searches for. I decided that it would be easier to manage such processes via a relational database, so on Friday and mapped out a structure for the playbill data and bean working on an import script that would process the JSON files. Lots more to do in the coming weeks!
Week Beginning 26th January 2026
Once again this was a week of many different projects. I spent a fair amount of time on Monday preparing a CV for a Leverhulme bid that Clara Cohen is putting together. I hadn’t worked on a CV for at least 12 years, so it took some time to look back through everything and prepare the text. I spent most of the next couple of days on the Burns Supper Map project, with the bulk of this time spent writing a specification document that describes the map I’ll create and the data it will use. It took quite some time to prepare the document – not just the actual writing of it, but thinking through how the data will be presented and how users will interact with it. I’d completed a first draft by the end of Tuesday and sent it to the team for feedback. They are also going to send it on to other interested parties and hopefully they’ll get back to me next week and I can begin work developing the site.
On Tuesday I also met my fellow College of Arts and Humanities developers for one of our coffee and catch-up sessions and it was a good opportunity to hear what they’ve been up to and discuss some of the technical issues we are all currently dealing with. On Tuesday I also met with Jennifer Smith to prepare for our Speak For Yerself talk in Edinburgh next Monday. We’re just about there with our preparations and hopefully all will go well.
On Wednesday I set up a bare-bones WordPress site for the Playbills project, as the subdomain and server space I’d requested had come through. For the moment this does not include Solr (which we’ll need for the searches) and IIIF (which we’ll need for the images), but I’m intending to start developing things locally on my laptop so we don’t actually need these things just yet anyway. I’ll need some input from the project PI Deven Parker on things like images to use, themes, fonts, logos and colour schemes before we can go live with the initial project website and there’s no real rush to do this. I’m hoping to start working with the project’s YAML files to extract things like a list of distinct venues next week.
I spent most of the rest of the week working on the Place-names of Armagh project, working with the existing data and creating scripts to import all of the existing data they’d sent me relating to placenames, historical forms and sources into the CMS. There are now 2999 sources in the system and 233 place-name records, connected to 3412 historical forms. Almost all historical forms connect through to a source (3408). There was an issue with a source with ID 203 that was referenced in the historical forms spreadsheet but no source exists with this ID. It took some time to write and test the import scripts, and it’s possible further tweaking will be required, but I’m pretty happy with how the process went.
I also added in the available grid references. These were not in the CSV data I’d been sent, but were included in the shapefile data that I was able to load into QGIS. I was able to export this data as a CSV file from QGIS, and thankfully the IDs in this file corresponded to those of places in the other CSV files I’d been sent, so I was able to join things up and import the grid references. 141 out of 233 placenames have grid references, but what I haven’t had time to do yet is to use this to populate latitude, longitude and altitude. This is something I’ll need to look into next week.
I also mapped the ‘Status’ codes onto the classification codes in my system. I’ve added some new classification codes taken from the new data (‘Ro’ for road system, ‘X’ for ex nomine, ‘M’ for minor place, ‘D’ for district). Other status codes have been mapped onto existing classification codes. ‘H’ has become ‘R’ (relief), ‘V’ has become ‘S’ (Settlement). I haven’t imported ‘DY’ as this didn’t seem to fit with the rest of the codes.
I also imported all Parish, Barony and Townland associations for each place. Some places have a different parish in the 1865 and 1961 columns and in such cases the 1865 parish is associated as a ‘former parish’. I also imported the map sheets.
On Friday I had a useful meeting with the Armagh team where I talked them through the data in the CMS and we discussed some of the issues that cropped up. I now have a list of updates that I’ll need to make to the data structures and the CMS, including trying to automatically extract Irish names and translations, renaming ‘Discovery’ maps to ‘1:50,000’, removing the separate language fields from the historical forms, ensuring Townlands can be selected from a list, as with parishes and baronies, and adding in a new ‘Previous suggested form’ Y/N field to the historical forms.
Also on Friday I made a couple of updates to the Anglo-Norman Dictionary to ensure that ‘M.E.’ appears as ‘English’ in the entry and search pages. I also made a couple of minor updates to the Dynamic Dialects site and helped sort out an access issue that an RA was having with one of the project websites.
Week Beginning 19th January 2026
I worked on many different projects this week, but the one I spent the most time on was the Dictionaries of the Scots Language. I’ve not done much work for the DSL since the intensive period I spent developing the new website interface and deploying it on our test server ahead of the face-to-face meeting in mid-November. I had a list of further updates I needed to make following on from this meeting, but I needed to work on other projects since then and hadn’t got around to it. I’d also received a number of emails about changes to the presentation of entries reflecting the structural changes to the entry XML that I’d put to one side.
On Wednesday I had an online call scheduled with the DSL team to discuss the new front-end and it seemed like a good opportunity to get back to grips with all of my outstanding DSL tasks. This mainly involved making updates to the XSLT on our test server to tweak the layout of various items in the new entry XML structure, such as adding commas between tags when they are rendered, ensuring certain tags or attributes that weren’t getting rendered before appeared in the generated HTML, updating the styles of certain elements like the content warning labels, fixing a few bugs such as the ‘sticky’ heading not displaying in certain circumstances. There were at least 20 such items that needed investigating, fixing and testing, so this took quite some time to work through, but I managed to complete it all during the course of the week.
The meeting itself was very useful and as always it was good to catch up with some of the other DSL team members. There’s going to be a big push towards getting the new website interface ready for publication this year, and I’m obviously going to be involved in this process. I already have a number of items I need to sort out with the new interface and I’ll try and get started on these over the coming weeks.
I also spent a bit of time this week working for the Anglo-Norman Dictionary, investigating a strange occurrence with the publication of updates to entries, which turned out to be a user rather than a system issue, reinstating the links out from entries to the DMF dictionary, as their website is now properly back online again, and tweaking the wording of the quick search and ‘jump to entry’ text throughout the site.
I also did small amounts of work for several other projects, such as updating the licensing statements across the Seeing Speech and Star sites, fixing an issue with the ‘download song’ facilities on the Editing Robert Burns site, sorting an issue with the HiMuJe Malabar site, submitting my expenses from the Zurich workshop, exporting some SCOSYA data for Jennifer Smith, helping to sort out an issue with the Helsinki Corpus, and having a conversation with Clara Cohen about a new proposal she’s putting together. I also made some further updates to the VARICS look-up system, adding in some introductory text, some further references, and reworking the measurement processing so that when a red or amber result is given new textual sections about what this means and what the next steps should be appear underneath in collapsible accordion sections.
Also this week I had a meeting with the Burns Supper Map team to discuss the data that is now coming in and how and when I should start working on a new interactive map to visualise it. I also put in a request for a new subdomain for the project that was set up by the end of the week. Next week I’ll probably write a brief specification document for the front-end.
My final project of the week was the Place-names of Armagh project, for which I started working with some existing place-name data for the area. There are around 230 place-names and several thousand historical forms and I spent quite some time researching how the data was structured and how it might be mapped onto the Glasgow place-names system. This included analysing the geospatial data, including shapefiles for Townlands and what I though was Parishes (but actually turned out to be the same as for Townlands). I had hoped to be able to import the data by the end of the week, but my analysis of the data raised a lot of questions that still need to be addressed, and I’ll need to continue with this next week.
Week Beginning 12th January 2026
This was my first proper week back at work, having spent most of last week travelling and attending a workshop in Zurich. I spent a bit of time working on the Bilingual Thesaurus of Everyday Life in Medieval England, looking into issues that had cropped up at the workshop. Someone had spotted that the start and end dates for some lexemes appeared to be the wrong way round and last week I discovered there were 197 such cases. I had an ongoing discussion with the project PI Louise Sylvester about this. She sent me a spreadsheet that contained updated data for the thesaurus, with the idea being that we could check the erroneous dates against this. However, the spreadsheet was created for a later project than the BTH and had both a different structure and different data. For example, some categories in the online BTH were not included and many categories in the spreadsheet featured different or larger numbers of lexemes. The dates were in a different format, featuring ‘ante’ and ‘circa’, plus a question mark to denote other uncertainty and a plus to denote continuation. The BTH features none of this – just start and end dates. The spreadsheet also featured no links out to the MED and the AND, only links to the OED. We did wonder whether we should replace the online BTH with the data from the spreadsheet but all of these issues mean this just wouldn’t work. Instead we decided that I would (at some point) write a script to identify lexemes in the spreadsheet that are not in the online BTH and we can see about incorporating them. In the meantime I fixed the 197 lexemes that had their dates the wrong way round.
Also for the BTH this week I implemented an option to order the lexemes in a chosen category alphabetically, by first attested date or length of attestation (within the AN or ME section), where previously all lexemes were ordered alphabetically within each section. This is something that was raised at the workshop, and something I wanted to implement as it’s a useful feature. I’d already included this option in the main HT and parts of the code for it were lurking in the BTH code in an inactive state, although I needed to rework this as the main HT handles dates in a more complex manner. The update required changes to the database, the CSS, the PHP and the JS scripts, but it’s all now live and the site remembers your choice during your session, so if you select ‘length of attestation’ in one category and then navigate to another this is remembered. Below is a screenshot showing a category with the lexemes ordered by length of attestation:
This week I met with Jennifer Smith to discuss the talk we’re giving about Speak For Yersel in Edinburgh in a couple of weeks. We had a good chat and made a plan about writing our respective sections. I then spent about a day preparing the slides and text for my section and sent everything over to Jennifer so she could work on her parts.
Also this week I did a little bit of work for the AND, updating links from AND entries to the DMF, as their site has changed, which broke all our links. I thought I’d found a way to link through to their corresponding entries but unfortunately their URLs now include a session variable that expires after a while, and the URL doesn’t work without a valid session. This means it’s not currently possible to link to their entries so for now I’ve had to remove the links. Apparently they are working to fix things so hopefully we’ll be able to reinstate the links at some point.
On Friday I met with Deven Parker to discuss her Playbills project and the requirements document I sent her before Christmas. We discussed a few issues that had been raised in the feedback on the document and made a plan for the coming weeks, during which I will begin to work with the data and will start developing the online resource.
Other tasks I tackled this week included replacing the data I’d uploaded for the VARICS project last week with a new version I’d been sent, and also making several tweaks to the code and content of the lookup feature. I also changed the language abbreviation ‘Ga’ to ‘Ir’ in the place-names of Armagh content management system and fixed a typo in the Hummell edition on the Burns website that went live before Christmas.
Week Beginning 20th October 2025
I was on holiday last week, although I kept up to date with my emails whilst I was away, so at least I didn’t have a backlog waiting for me when I got back. I did have quite a lot to do this week, though, and the most pressing was to complete the rollout of the new XML entry structure for the Dictionaries of the Scots Language. I’d begun this process the week before my holiday, but had run into difficulties getting the Solr indexes set up on the server, and while I attempted to resolve this with Andy in IT Services we ran into difficulties were were unable to resolve before I headed off.
After further investigation and discussions with Andy this week it turned out that there were two issues. Firstly the ‘conf’ directory for Solr indexes needed to go into a different directory on the server compared to how things work on my laptop, and how things used to work on the server. Rather than appearing in the index’s own directory, the ‘conf’ directories need to be stored in directories that have the index name within the configsets directory, otherwise the creation of the index fails. Secondly, when Andy did manage to successfully create the first index the week before last he unfortunately got the name slightly wrong, which neither of us spotted in the rush to get things finished on the Friday afternoon before my holiday. This then prevented the command to index content from working, as Solr was attempting to post files to an index that didn’t actually exist.
Thankfully these issues have now been sorted, the ‘entries’ and ‘quotations’ indexes now work. I was then able to update our test website to use the new indexes, meaning the website now includes the display of the new entry structure, data extracted from the new XML structure and also the map popups. The DSL team can now experiment with all of this and get back to me with any tweaks or updates that they require.
Also for the DSL this week, I dealt with the last item on my list of feedback for the map of regions and dialect areas. I’d suggested that we should include facilities to cite or share views of the map, and also make use of the URL shortener I’d developed for the place-names resources and this was something the team wanted me to implement. By the end of the week I’d completed a first version of this and had added it to the map.
Initially I had intended to have a ‘cite’ popup as I use on the place-name resources, but then I thought it would be better to use the same ‘share’ options that we use elsewhere on the DSL site for consistency. I added the ‘share’ buttons in a row at the bottom of the left-hand menu and got things working with the full page URLs appearing in the ‘share’ output.
However, I ran into something of a brick wall when implementing the URL shortener. I tried many different approaches to get the shortened URLs to replace the default page URLs in the ‘share’ options but nothing would work. The reason is that the URL shortener posts data to the server to generate the shortened code and then saves this with the full URL in the database, but despite the ‘share’ plugin allowing custom URLs to be added in, there does not appear to be any way to make its code hang around until my script connects to the server and generates the shortened URL – it just jumps ahead, thinks no custom URL exists and uses the default page URL regardless.
After attempted many different ways to get around this, but failing to find anything that would work, I then decided to revert back to my ‘cite’ popup option, as this would enable me to generate the shortened URL when the popup is opened, meaning it’s ready and waiting when the user clicks the ‘share’ option contained within the popup. Plus in addition to the ‘share’ options we can also include the citation styles for people who want them. Thankfully this approach has worked, as the screenshot below demonstrates:
You can see the shortened URL in the ‘cite’ options (redacted as the resource is not yet live), and this shortened version is also used in the ‘Share’ options. This is still just a first version of the ‘cite / share’ feature, however, and further changes may be required. For now there is no custom text in the ‘share’ or ‘site’ options describing what the map shows. It would be quite a lot of work to add this in, so I’ll just see if the DSL team want me to do it. It would mean that rather than just saying ‘Map of regions and dialect areas’ a citation might say something like ‘Map of the regions of Scotland centred on [lat,lon here] with regions Ayrshire, North Ayrshire and West of Scotland selected’.
Also this week I had a meeting with the HiMuJe Malabar team to discuss the development of the interactive map. It was a very useful meeting and we have a much clearer idea of how things will develop now, and how the map will relate to the digital edition, which is something that is being managed at a partner institution. This led on to discussions about issues Ophira was having with displaying the digital edition files in Oxygen on her PC. The digital edition is one long block of prose and this is causing Oxygen to get rather laggy when editing it in the ‘author’ view. I asked my colleague Luca to have a look at this, as he has more experience with Oxygen than I do, and he was able to come up with a solution that splits the single XML file into multiple smaller files. This approach seems to have worked very well, but we need to see what the project partners make of this change before it can be adopted by the project.
I also met with Henry Ivry this week to discuss the development of a new online resource for the Beniba Centre. We had a good chat about the various options and what might need to be done and Henry is going to get back to me with some further information that I can then use to get things started. I also made a bit of progress with the place-names of Armagh project with Mícheál Ó Mainnín at Queens University Belfast. We now have a request submitted for the domain and hosting and hopefully this will all get set up in the next couple of weeks.
I also fixed an issue with the display of an entry in the Anglo-Norman Dictionary. A malformed <link_loc> was causing issues with the page. The <link_loc> should have a comma (1,175c), but this was missing. The JavaScript that processes the references expects to find a comma and when it doesn’t it throws an error, meaning all of the other JavaScript stops working (including the parts that rearrange the author’s initials and numbers). I offered to update the code to deal with this a little more gracefully, but the editor Geert didn’t think it was an issue that would crop up very often and didn’t think it was worth doing.
Also this week I was contacted by Alison Wiggins to ask me to do some work on her Mary Queen of Scots Letters project, which has lain dormant for several years. Alison now has some time to work with the data and get things moving again and she had a few questions and requests, which I attended to.







