Category: Misc place-names
Week Beginning 10th August 2026
I was back in Glasgow and back to a five-day working week this week, as the summer holiday period drew to a close. I spent a fair amount of time this week continuing to work on the DOST Auld Laws project. Luca has been working on the XML queries I’d specified for the advanced search and I was able to test them out and give feedback on them. These are mostly all working as I’d hoped, which is really great, and I was able to begin working on the front-end aspects of the advanced search, with the aim of connecting all of this into the queries Luca had created.
However, I also had to spend quite a bit of time working with the source files, as the project PI Joanna Kopaczyk-McPherson had realised that two of the eight documents should really be split up into smaller sections. This was no straightforward task, as not only did it require the XML documents to be split into smaller sections, but many other aspects needed to be updated. The document IDs needed to be changed, which mean IDs used for image filenames throughout the documents also needed to be changed, with the filenames of the actual images also then needing to be updated too. For page navigation in the site data is stored in a database and this also needed to be updated, and the XML files stored in eXist for search purposes also needed to be updated. There were twelve steps I needed to follow for each required split, which took some time, but thankfully the process went pretty smoothly and we ended up with 13 documents in the site instead of the original 8.
There were further issues to come, as Joanna has now noticed that the documents need further edits and tweaks, and not just minor changes to text but structural issues such as the insertion of omitted lines. This is going to be very difficult to do as lines are linked to coordinates in the images, which was all handled via the Transkribus tool. We exported the XML files from Transkribus month ago and many major changes have been made to the files since the export so it’s not going to be possible to re-import them. Manually creating new lines in the XML files as they are now will not have the connections through to coordinates in the corresponding image files, so we’ll end up with inconsistent data. It’s not a great situation to be in, and ideally all editing of the documents should have been completed in Transkribus before we exported the files, something I’d mentioned back when I undertook the export process months ago. We haven’t reached a decision on how best to handle this yet and we’ll continue to discuss the options next week.
Despite all of this I did manage to work on the advanced search, creating the advanced search form which, as specified, features a textbox where you can enter some text, a list of documents from which you have the option of selecting the ones you’re interested in and a list of tags that you can either include, exclude or limit your search to, as you can see in the following screenshot:
As of yet I have not implemented the tag limits, but the limit by documents is operational. This connects through to Luca’s new eXist-db queries to perform a search limited by documents and as with the quick search, you can also sort the results by words to the left and right of the term, although I’ll need to get Luca to look into how the KWIC is generated for the advanced search results as they don’t seem to cross line boundaries, unlike the quick search. The advanced search results display information about the documents you’ve selected and if you press on a search result to load the corresponding page the link back to the search results takes you to the right place. There’s still lots to do. The limit by tags is the biggest thing and will probably take quite some time. I also still need to add in an option to cite a specific search result page and add in an option to refine your search, which will remember the options you previously filled in when you return to the search form. An option to clear the search also needs to be added. I’ll continue with this next week.
Also this week I participated in two meetings about the place-names AHRC proposal I’m involved with, and this is coming together very well. We now have an outline proposal completed and pretty much ready for submission.
I also spent a bit of time continuing with the travel routes for the HiMuJe Malabar project, adding in a few more travel routes that had been prepared and made a few tweaks to the XSLT that generates the entry HTML for the new DSL interface, fixing some issues with the layout of the new ‘combs’ sections.
Week Beginning 3rd August 2026
I worked a total of four days over the past two weeks, and was on holiday for the remainder. During this time I had a meeting to further discuss a place-names related AHRC proposal I’m involved with. I can’t really say much more about it at this stage, but the proposal is coming together. I also had to spend some time working with IT Support and Luca to figure out why our local server kept going offline repeatedly. It looks like this was caused by the server getting swamped by requests from one particular source (almost certainly bot or AI) and thankfully IT Support were able to block this, after which the server was stable again. It’s something we’re going to have to keep looking out for in future. I also spent a bit of time working with Luca to get automatic WordPress updates working on our local server, as the way sites had been set up meant that the setting was not working. Luca managed to find a solution to this, which is really great.
I also spent a bit more time working on the Burns Supper Map, creating a record for it on this site (see https://digital-humanities.glasgow.ac.uk/project/?id=156), adding more suppers that had been submitted via the survey and making some requested edits to existing suppers. I also set up access to Google Analytics for the other two members of the project team.
In addition, I investigated an issue with the Scots Syntax Atlas after someone suggested that the linguists’ atlas was looking somewhat blurry. I managed to figure out why this might be the case, although I’m not entirely sure whether this is a new issue or if the markers always looked like that. I contacted the project PI and suggested a couple of updates, but I haven’t heard back yet so will need to wait and see what she says.
I spent most of the remainder of my time working on updates to our new test interface for Dictionaries of the Scots Language and working through the list of outstanding items for the DOST Auld Laws project. For the DSL I completed the updates to the bibliography page that I began working on a couple of weeks ago. I implemented pagination of the entries associated with bibliographical items, with navigation bars appearing above and below the entries, with 20 appearing per page and ‘jump to page’ buttons also appearing, just like with the search results. This works pretty well, but is somewhat cumbersome for someone like Sir Walter Scott, who is referenced in 2861 entries, split over 144 pages. We don’t have this issue with the search results are these are capped at 500 (25 pages) so we might need to think of other ways of handling this.
I also ensured that headword searches that don’t yield any results automatically perform a fulltext search for the term supplied. This works on the live site, but only when both dictionaries are selected. With the new site there are separate quick searches for SND and DOST so the additional search wasn’t being triggered. It is now, as is the advanced headword search for both dictionaries.
I also made tweaks to the DSL’s new regional map based on feedback I’d received – adding in some content where we previously had placeholder text and ensuring the ‘About’ popup didn’t disappear off the bottom of smaller screens and a few other small updates. I then began to look at the ancillary pages and how we can make them look a bit nicer. I spent a bit of time on the ‘Word of the Week’ page and liaised with William Ashford, who is responsible for such content about this, and further updates that were going to make to the ancillary content closer to the launch date of the new site (which will hopefully be in November).
For the DOST Auld Laws project I added the top navigation bar that will link the site in with the SCOTS Corpus and CMSW. I added in copyright information and added facilities to download page images and the XML files for each document. These being up a pop-up asking for users to abide by the license before leading to the actual content, which hopefully won’t be too annoying. I also added in the ‘cite’ popup to all document pages, which took a little time to implement, and added in Google Analytics. I also made the image thumbnails on the document overview pages smaller and placed them in a collapsible section that is closed by default, plus I removed the introduction to the documents page, as this will be covered by the homepage.
I also removed some pages that didn’t have content (e.g. blank pages) from the beginning and end of some of the documents and I added a feature to turn off and on the line highlighting feature. The highlighting feature allows the user to click on a line of text in the image or text for a page and for that line to be highlighted in both the image and the text, which is pretty nice. Unfortunately the line highlighting gets in the way of the image viewer’s zoom and pan functionality on touchscreens, making it a somewhat unreliable and frustrating experience. This new feature removes the option to ‘click’ on a line, meaning pointer events are not intercepted and make their way reliably through to the image viewer, which works much more smoothly.
Also this week I had an email conversation about user feedback and walkthough videos for the STAR resources and booked my accommodation for the DHC conference in Sheffield. Next week I’m back in Glasgow and back working a full week, with summer holidays all over.
Week Beginning 13th July 2026
This was a four-day week for me as I’d taken Friday off (and I will also be off for the first three days of next week). I finally managed to assign some time this week to implementing updates to the new Dictionaries of the Scots Language interface based on feedback that had been sent to me earlier this year. I spent most of Monday and Tuesday working on this.
I can’t share any screenshots of the new interface at this stage, but I updated the search results box in the entry page to remove the tab for the other dictionary when performing a search (quick or advanced) for a specific dictionary. This avoids misleading people as it otherwise the tab displays zero results for the other dictionary when in fact it just means the dictionary wasn’t actually searched. Now when a quick search is performed, the other tab heading is replaced by a link to search the other dictionary. Pressing on this performs whatever search you’ve executed (quick or advanced) on the other dictionary.
I also tweaked the font colour of the inactive tabs. I realised that the white text on grey made it look like the tabs were disabled, when they’re not, they’re just inactive. I therefore made the font darker, which I think works a lot better. I then added in a button that scrolls the page to the search / browse box. This appears above the entry header (and in the ‘sticky’ header that appears as you scroll the page) and only appears on narrower screens (where the infobox appears below rather than beside the entry text). Where the entry is in the search results the text is ‘Scroll to results list’ with a down arrow. Otherwise the text is ‘Scroll to browse list’. The DSL team had requested that results term highlighting should be off by default, so I made this change too.
I then began to rework the bibliography page based on feedback. We’ve decided to go with the version of the bibliography page that displays the quotations from any associated entries in addition to the headwords and links through to the entry pages. This required some reworking of the API so that rather than returning individual citations, it brings back entries with each associated citation as part of this. This means that multiple citations for an entry no longer appear as separate items in the list but are grouped by their entry, much like the quotation search results. I also updated the count above the citations to display both the number of entries the item appears in as well as the number of citations (e.g. ‘Cited 718 times in 597 entries’) and each citation also includes its date as well now. However, the order of citations needs to be the order they appear in the entry and not date order, otherwise the links through from the citation may end up taking you to the wrong one in the entry (as I discovered when I set things to date order).
The display of entries and citations is not exactly identical to the search results. There is no sparkline as unfortunately this is not included as part of the bibliography data and I’d need to rework the database and API in order to include them. Also, the quotations are in a larger font than the ones in the quotations search results as they seemed a bit small, and the headword isn’t highlighted in the quotations as this is something performed by the Solr search engine and isn’t available as things currently stand for the bibliographies. I still need to add in pagination, which I didn’t have time to work on this week but will hopefully implement soon.
I also participated in a Teams meeting with the DSL this week that involved their interns reporting back about user engagement and observations about the website. It was very interesting to hear their feedback and it will give us lots to think about as we continue to improve the resource.
Other than working for the DSL, I also made some last-minute updates to the Burns Supper Map before its official launch over the weekend. The resource is now available for anyone to use at https://burns-supper-map.gla.ac.uk. I wasn’t able to attend the launch as I was away on holiday but from what I’ve heard it was a great success.
I also spent a bit of time fixing an issue with the Anglo-Norman Dictionary where certain entries that had an apostrophe in their headwords (e.g. j’) were not loading while others were. The reason for the discrepancy was because some headwords had curly apostrophes and other had straight ones. The straight ones were getting encoded (e.g. “j'”) and were then not getting found. Once I managed to figure this out I was able to fix the issue.
On Wednesday I had a lengthy online meeting for the HiMuJe Malabar project to discuss the interactive travel routes. The meeting lasted about two and a half hours but it was worthwhile as we all have a much clearer idea of how to proceed with the routes now. I just need to wait until the team sends me some initial travel routes using the spreadsheet template I sent them and then I’ll be able to continue my work on this.
On Thursday I participated in a Teams call with colleagues from Nottingham and Cardiff Universities to discuss a place-names proposal that I am likely to be involved with. I can’t really say much more about it at the moment, but it’s all sounding very interesting. I spent a few hours after the meeting writing a document containing some initial thoughts about how the technical infrastructure for the project could work.
I was then off on holiday on Friday and I won’t be back at work again until next Thursday.
Week Beginning 11th May 2026
This was a pretty busy week that had me dividing my time between five main projects. I spent most of Monday working on the Burns Suppers project, beginning development of the ‘facts and figures’ popup. I added a button in the ‘Home’ menu labelled ‘Facts & figures’, that features a pie chart icon. Pressing on this opens the facts and figures popup part of which you can see in the screenshot below:
The popup features the narrative summary section, with the figures in this section being dynamically generated. It took quite some time to write the code to generate them, but at least there will be no further work to do when we import more data. For now all percentages are the percentage of the total number of suppers rather than a percentage of the number of suppers that supplied the data, so for example, the ‘annual supper’ figure is 39% rather than displaying 83% (based on there being 420 suppers that actually have frequency data). This is because I realised that without a lot of additional explanatory text users will likely think our figures are wrong. If people only see the total number of suppers (886) and the number that are annual (347) then displaying 83% will be misleading. We’d have to include lots of additional data such as “of the 420 suppers that included frequency data, 347 (83%) were annual events”. Depending on feedback I may implement this later.
I’ve only implemented one graph in the pop-up so far, which is the graph of countries. This is a long list, but I think it works ok. I also included an option to switch from a graph ordered by number of suppers to an alphabetical version, and this is all operational. I’ll continue to add further graphs next week.
I spent most of Tuesday working on the interactive map for the HiMuJe Malaber project. I’ve created an initial version of the map now, which uses my map menu interface that I’ve used on several other resources, and below is a screenshot:
By default the map uses the ‘Watercolour’ base map, as we used for a static map on the main site (https://himuje-malabar.glasgow.ac.uk/about/summary/). You can also switch to a satellite map using the ‘Change the base map’ buttons. Currently markers as just displayed as red dots, and if you hover over them the ‘Preferred Name’ from the JSON file is displayed as a tooltip (although I’ve noticed that some places don’t seem to have a ‘Preferred Name’). There are no popups or filters or anything like that yet. Also note that any places in the JSON file that don’t currently have latitude and longitude values are not displayed as there is nowhere to ‘pin’ them.
You can zoom and pan the map as with Google maps, and the icon in the bottom right makes the map full screen. The map also works on mobile devices (hiding the left-hand menu using the ‘<’ button above the menu helps when using a mobile device). Other than the base map selection options, nothing works in the left-hand menu yet. ‘Places’ will eventually include the place filters. ‘Travel Routes’ will list travel routes involving people and organisations once this data is available. There’s still a lot to do, and I’ll hopefully begin work on the place popups and the categorisation options next week.
I spent most of Wednesday working on the Place-names of Armagh project. I’ve updated my ‘Irish form / translation’ script to remove ‘More details: Unverified’ and to also attempt to split the translation up based on apostrophes. This has been a bit of a nightmare, firstly because apostrophes are not only used to denote the translation but appear within the text, and secondly because what looks like an apostrophe can actually be many different characters, including straight, curly opening and closing apostrophes and several HTML codes that are rendered as apostrophes but are stored as codes. This has all made splitting the text up rather challenging. However, I’ve got something that mostly works.
The script now deals with things like Ir. <em>Coill Uí Fhloinn ‘O’Flynn’s </em>wood’ and outputs “O’Flynn’s wood” as the translation. Where there are multiple sections some unnecessary apostrophes appear, for example: “Ir. <em>Tír Garbh </em>’rough land or district’ or perhaps Ir. <em>Baile Uí Aodha </em>'<em>O’Hugh’s</em> homestead or townland’” results in: “rough land or district’ or perhaps Ir. Baile Uí Aodha ‘O’Hugh’s homestead or townland” This will probably need some manual fixing. The script also now processes rows that don’t have ‘Ir.’ and italics, for example “Críonchoill ‘withered decayed wood’” now has Irish form “<em>Críonchoill </em>” and translation “withered decayed wood”.
I sent an Excel version of the script output to the project team as this will likely require some (but hopefully not too much) manual intervention to fully sort out. For example, the text “Eng./Sc. ‘hill of the military camp’ or perhaps Ir. <em>Mullach na Críne </em>’hilltop of decay’ or Ir. <em>Mullach na Craoibhe </em>’hilltop of branch, tree’” will need some work as the translation omits the ‘Eng./Sc.’ Text as it’s before the first apostrophe.
The other thing I’ve managed to do today is to set up an initial version of the public map interface. This currently takes quite a long time to load as it’s processing a lot of data. As with the other sites, I’ll create a cached version of the data once we’re ready to launch, which will be much faster to load. The reason it’s not in place now is that a new version of the cache will need to be generated any time you want subsequent changes made to the CMS to appear on the map, and we’re still very much working on the data.
For now, three base maps are available (satellite, satellite with labels and relief). We can add more in later. There are also no parish boundaries or townland / barony boundaries as I don’t have this data yet. I also still need to work on the colours of the markers as there are no colours for a lot of the classifications so they’re defaulting to purple. There’s also nothing in the elements glossary as we don’t have this data yet. But the search, browse and categorisation options all work. For example, here are the placenames beginning with ‘A’ categorised by altitude:
There are some issues with the data. For some reason there are three place-names miles away from Armagh, around Ballybofey, and there are a couple of place-names appearing in the Irish Sea. We also have an issue of the same coordinates being used for multiple places, thus resulting in markers appearing on top of markers. There’s also an issue with the accuracy of markers too. For example, the marker for ‘Lowry’s Lough’ is found about 500m south of the actual body of water. But the good thing about having this map available (despite the speed issues for now) is that it will help when working on the data.
I spent all of Thursday and most of Friday morning learning how to use the Exist-DB XML database that I’m hoping to use for the DOST Auld Laws project and developing the query that will eventually power the quick search, and form the basis for the advanced search. I installed Exist on my laptop and managed to set up a collection for the project’s XML files, which I then uploaded into the system. I followed the documentation available on the Exist website and was able to create a full-text index for the collection, and I followed a useful tutorial here: https://dh.obdurodon.org/php-xquery.xhtml about how to query Exist using the REST interface. Setting up and querying the texts in Exist was all new to me and there was a lot to try and take in, with many configuration options that were not all that easy to follow in the Exist documentation. I ended up using ChatGPT quite a lot to help me figure out how everything should work and why some of my initial tests were not producing any results. This proved to be extremely useful and really helped increase my understanding of Xquery. I was able to get a search working that queries the full text and returns contextual snippets for each result, together with the IDs of the line, page and document, which is everything I’ll need for the quick search. On Friday I asked Luca to help set up the necessary collection on the server and hopefully I’ll be able to get an initial version of the quick search working on the website next week.
This left me with a few hours on Friday afternoon to devote to the redevelopment of the Mapping Metaphor resource, for which I’m creating a unified view of the data, joining both the English and Old English datasets together. I used this time to implement the visualisation card view when viewing connections between specific categories. The combined card view shows a bidirectional arrow if the OE and E directions differ. It also defaults to the E strength. The counts of lexemes in each category include both full and OE counts and the examples of metaphor feature both OE and E examples, with OE coming first. If the metaphor exists in the OE data then the ‘start era’ now defaults to OE (but as of yet I’ve not added in a further highlight in the timeline to show when the first non-OE occurrence was documented). Below is an example of the combined card view:
In addition to the above I also looked into an issue raised by Ann Fergusson for the Dictionaries of the Scots Language regarding accented characters, search results and entry slugs. This took some time to investigate but I think my response proved useful. I also had email discussions with Deven Parker about her Playbills project, which she now has time to look into again. I’ll probably be working on this again next week.
Week Beginning 4th May 2026
This was a four-day week as Monday was a bank holiday. I spent most of Tuesday and Wednesday continuing to work on the DOST Auld Laws project, working with the XML files. The XML files generated by Transkribus and exported by the tool as TEI XML contained many elements that were not valid TEI elements, such as the custom <Aitken> element that had been applied to notes added by A J Aitken. My first task of the week was to write and apply transformations to the XML files to convert them into fully valid TEI. I achieved this using XSLT, which is a language I find very unintuitive and frustrating to work with, no doubt exacerbated by the fact that I don’t work with it very often. I struggled to get any transformations to run initially, and ended up turning to AI to figure out why my scripts were not working. In this instance AI proved to be extremely useful as it identified what the problems were (e.g. I hadn’t declared the correct namespace or used it when writing the rules) and really helped me to understand how everything fitted together. I still wrote the bulk of the code myself, but AI was very helpful in identifying errors or issues. By the end of Tuesday I had generated (and checked) a collection of TEI files that successfully validated in Oxygen, which was a good milestone to reach.
On Wednesday I then worked on the front-end for the project, figuring out how to transform the valid TEI XML into HTML for display in the ‘text and image’ and ‘text only’ views of document pages. As this was once more using XSLT I enlisted the help of AI to figure out specific issues that I was unfamiliar with. The biggest of these was how to pick out and process one single page from a document’s XML file based on the <pb/> element. I had no idea how to achieve this, and despite this being a fairly fundamental issue when processing TEI documents I didn’t manage to find any useful information online. However, AI came up with a solution (and equally importantly an explanation) in a few seconds and I was then able to incorporate this into my code, transforming the contents of one page, whose ID was passed to the XSLT script as a parameter, to HTML with a variety of styles applied to the various elements.
I also updated the ‘click on a line in the image’ feature, and it’s now possible to deselect the line if it’s already selected – previously once you’d highlighted a line you couldn’t get rid of the highlighting, only move it to a different line but now if you press on the highlight it’s removed. The lines are also now connected to the text view – pressing on a line in the image also highlights the corresponding line in the text. You can also press on a line in the text to highlight it and the corresponding line in the image.
I also updated the height of the text pane so that it matches the height of the image pane and if the text is longer the pane scrolls. This ensures that if the text is very long it’s still possible to see the image, rather than having the entire page scrolling, which may result in the image not being visible when you’re at the end of the text. It is how we did things in Books and Borrowing, but having a scrollbar in a section of the page in addition to the browser’s scrollbar can be annoying for some people so I might revert to the previous layout depending on feedback. Here’s a screenshot showing the transformed text, a highlighted line and some of the formatting that’s been added:
Also this week I continued to work on the Place-names of Armagh project. I wrote and executed a script that generated Irish grid references and ITM values for all places based on their latitude and longitude, a task that I completed using the ‘GridRefUtils’ scripts as detailed here: https://www.howtocreate.co.uk/php/gridrefapi.php. I’ve used these scripts before on previous place-name projects and they’ve been hugely useful. I then ran a further script to generated altitude for the place-names by connecting to the Google Maps API, so we now have complete geospatial data for all of the places (other than the 386 that didn’t include Easting and Northing data in the original spreadsheet). I also added in a new parish and barony in a different county that one place-name requires, and had discussions with the team about the splitting of analysis data into Irish forms and translations. Removing the Irish forms from the translation field is going to be rather tricky to automate as the Irish forms often form part of the translation. It’s looking like I’ll need to automate the transformation of some of these, with the rest then requiring manual intervention.
I also continued working on the redevelopment of the Mapping Metaphor resource to create a combined English and Old English map, something I began last week. This week I managed to get the combined view of the drilldown of the visualisation working. The counts represented by the yellow circles also use the combined data, and these are also displayed in the pop-up card, for example, if you select 1K02 Creation and press on the yellow line or circle for 1B Life. In the combined map there are connections to 6 categories in Life whereas there are 5 in the E map and 2 in the OE map (one of the OE ones is also present in the E map, which is why the combined total is 6). The combined pop-up card also now lists the number of OE lexemes in addition to the E lexemes, for example “1K02 Creation 1088 lexemes / 127 OE lexemes”.
I still need to implement the combined visualisation view of the search results, and also the card pop-up between individual categories, which is going to take quite some reworking (possible direction changes, combined example lexemes, timeline updates etc). I’ll hopefully find some time to continue with this next week.
Also this week I processed and added some more images to a few Burns Suppers, had a chat with Garrick Allen about a research project he’s wanting me to be involved with later in the year, and contacted Lindsay Balfour about a proposal she’s writing that will have some technical requirements. I also made a couple of tweaks to the Thesaurus of Old English website after Fraser go in touch with some suggestions, made some tweaks to the survey popup on the Dictionaries of the Scots Language website, gave some mapping advice to Renu of the HiMuJe Malabar project, and created an alternative version of one of the Anglo-Norman Dictionary’s textbase documents that strips out all Latin text.
Week Beginning 27th April 2026
I had several meetings this week, the first being to discuss the development of the map for Ophira Gamliel’s HiMuJe Malabar project. In the run-up to the meeting it became clear that no-one had read the specification document I’d written regarding the development of the map, which I’d sent out several weeks ago, but thankfully after I raised this it was distributed before the meeting and we had a very useful discussion about the map and how to prepare the data for the map. It feels like some real progress has been made in our joint understanding of the data and what needs to be presented, and following the meeting I sent out a proposed structure for the JSON data that will be generated from the digital edition files and a proposal for how we would document travel routes on the map. I also met with project RA Renu a couple of days later to discuss geoJSON data and the creation of polygons for certain locations in the data.
My second meeting was with the Place-names of Armagh people, and this was also very useful. We discussed some additional explanatory text found in the original data that I hadn’t spotted before, the splitting up of the existing ‘analysis’ data into Irish forms and translations and the creation of the front-end. Now all I need to do is actually find the time to work on all of this. My third meeting was with the Burns Supper team, at which we went over the ‘facts and figures’ popup that I will be developing. We also discussed the launch of the resource and a demo that was given on Monday and was very well received.
I spent a fair amount of time this week working on the DOST Auld Laws project, beginning with an analysis of the XML files in order to work out what non-standard elements generated by Transkribus needed to be converted into which TEI elements. It took some time to go through everything but it was very useful to get it all documented and subsequent email discussions with Joanna and Pia were really helpful.
I then ran a script to update the XML files to change the filenames, add in IDs and replace the referenced image filenames (both in the XML files and the image files themselves) with a more rational naming structure as I’d previously discussed with the team. For example, the file with the title ‘115 – Peebles B. Rec.’ was given the ID ‘peebles-b-rec’ and image files and references were renamed from (for example) ‘0001_115 001.jpg’ to ‘peebles-b-rec_001.jpg’. With this in place I then arranged with Luca for the images to be uploaded to the IIIF server, after which I could begin working on the front-end on the server, as opposed to on my laptop as I’d been doing previously.
As part of this work I also decided to migrate the document data I’d extracted from the XML from JSON to a MySQL database. The reason for this was that some of the document JSON files were rather large (a couple more than 6MB) and it seemed rather inefficient to load all of this data in just to work out counts of the number of pages and such things. With all of the data in a relational database it’s much easier to query and return just the data that is needed.
During the week I completed work on an initial version of the image parts of the site. The user interface is by no means complete and is fairly rudimentary – we will decide thinks like fonts, colour schemes, illustrative images and ancillary content later on. For now there is just placeholder text on all but the ‘Documents’ pages and the search option doesn’t work yet. However, it is possible to browse the documents and all facsimile images contained in each. Pressing on the ‘Documents’ menu item brings up a sub-menu listing the documents, plus an ‘overview’ page which is currently empty. If you select a document you’re presented with an overview page. This can feature a description of the document (there’s just placeholder text for now) a link to open the document at the first page and a randomly selected image displayed to the right of the description for illustrative purposes.
Beneath this is a section through which you can see a count of the number of pages and access a list of thumbnails of every page. For some of the longer documents it can take some time to load in all of these thumbnails. We can maybe have this section hidden by default, or possibly paginated. Pressing on the ‘open document at first page’, the random image or any of the image thumbnails opens the page in question, a screenshot of which is shown below:
I’ve borrowed much of the layout of this page from the work I did on the Books and Borrowing project. The page features a navigation bar at the top and bottom, through which you can navigate to the next or previous pages, or use the ‘jump to’ feature to select a specific page. There are three views of the page: an image and text view which displays both side by side and then text or image only views. When you select a view it is remembered as you use the navigation options.
As of yet there is no content in the text view (processing the XML will be my next task) so for now it’s all about the images. You can zoom and pan the image or open it full screen. You can also press on a line to highlight it. Eventually this will allow highlighting of the text from the image and vice-versa, but for now all that happens is the section of the image is highlighted. As mentioned in an earlier post, I adapted an existing ‘simplify coordinates’ script (https://github.com/dariok/page2tei/blob/master/simplify-coordinates.xsl) to generate simple rectangular sections for each Transkribus polygon, but we may need to have more complex shapes as in some pages (particularly the handwritten ones) the boxes are not very accurate. I’ve also spotted an issue on touchscreens (on my Android phone using Firefox, at least) whereby the boxes for the lines stop the ‘pinch to zoom’ feature working, meaning zoom will only work using the icons in the top left. I’ll need to investigate this further.
My next task will be to work with the XML, firstly replacing the Transkribus tags with valid TEI ones and then generating the text view for the website, which will link each line to the image. Once this is in place I’ll begin to think about the search facilities.
On Friday I spent a bit more time working on Sara Ponz-Sanz’s AHRC application and spent the rest of the day beginning work on the redevelopment of the Mapping Metaphor website. It took a while to refamiliarize myself with the code and to work out how things function, but once I made some progress with this I decided to start work on a new ‘combined’ version of the map. This will display the combined E and OE data but will not initially allow you to switch the display to just E or OE. There is a LOT of work that needs to be done to get this working fully, and I only made a start on things this week. For now only the visualisation works, and only the top level view of this works.
Working out the totals for the visualisation is not as simple as just getting the E and OE data and adding them together, as this would duplicate data for metaphors that appear in both sets. Instead, the new version of the visualisation grabs all of the E data and then only adds OE data for connections that are not already present. So for example, in the new visualisation pressing on ‘1K’ displays a connection to ‘3E’, which is not present in the E map but is present in the OE map.
In the combined map with strength set to ‘Strong’ and with ‘1K’ selected, pressing on the yellow line or circle for ‘1B’ the card shows 18 connections. If you do the same for the E map you’ll also see 18 connections, while the OE map shows 3 connections. The reason the combined map shows 18 rather than 21 connections is because the three OE connections already exist in the E data. Setting the strength to ‘Both’ in the combined map shows 38 connections between these categories, whereas there are 35 in the E map and 10 in the OE map, and this is because there are three additional weak connections between categories in ‘1K’ and ‘1B’ in OE that are not found in E. There is going to be a lot of figuring out how exactly to amalgamate the data across all views and all levels, which is going to be a pretty large undertaking. But I feel that I’ve begun to figure it out, at least.
Week Beginning 20th April 2026
I spent much of Monday and Tuesday this week working on the DOST Auld Laws project, which I’d begun work on last week. I set up a test instance of the Cantaloupe IIIF image server on my laptop and began experimenting with the OpenSeadragon IIIF image viewer. It was pretty straightforward to set this up and to create a test page where the viewer connected to an image stored in the IIIF server and enables the user to zoom and pan around the image. The trickier issue was setting up the infrastructure so as to allow lines of text in the image to be clicked on and highlighted, and for IDs associated with each of these line regions to be accessible by JavaScript code (to eventually enable the corresponding line in the text view to also be highlighted).
Luca is currently working on a digital edition that allows regions of an image to be highlighted when buttons are pressed on and he very helpfully gave me access to the test site he’s working on so I could see how things work. I had been expecting that the region data would need to be somehow stored in the IIIF manifests for each image and that the IIIF server would be a lot of the processing to enable the display of and interaction with regions in the image, but both Luca’s site and another one I was referencing stored the region data as coordinates pulled into the front-end from a separate source, such as a JSON file. I was very happy to go with this approach, as firstly it meant I didn’t need to spend ages investigating IIIF manifests and how to update them and secondly it means the region data is not tied into a specific server technology, meaning it will be easier to repurpose it in future if that technology changes.
The line data exported from Transkribus consisted of detailed coordinates such as:
1309,2018 1469,2068 1716,2037 1876,2123 2092,2037 2234,2111 2389,2062 2555,2123 2734,2062 2796,2111 2975,2068 3555,2130 3629,2080 3808,2111 4062,2037 4457,2043 4524,2080 4599,2018 5055,2099 5142,2049 5321,2099 5438,2043 5586,2086 5845,2043 5901,2080 6271,2062 6271,1901 6129,1944 5802,1839 5710,1920 5284,1889 5234,1938 5166,1883 4771,1907 4481,1796 4333,1907 4222,1870 4043,1957 3950,1907 3808,1920 3753,1864 3574,1913 3506,1852 3426,1907 3006,1839 2895,1895 2438,1870 2265,1913 2055,1815 1975,1852 1759,1839 1660,1920 1506,1920 1309,184
But for the most part these refer to simple rectangles and with our data this level of detail isn’t really needed. Thankfully there is a script available to simplify the coordinates (https://github.com/dariok/page2tei/blob/master/simplify-coordinates.xsl) and I was able to adapt this to export the pixel coordinates of each corner of the line rectangle, such as (for the above polygon):
1309,1796 6271, 1796 6271, 2130 1309, 2130
OpenSeadragon doesn’t work directly with pixel dimensions, but has a function to convert these into its required format (see https://openseadragon.github.io/docs/OpenSeadragon.Viewport.html#imageToViewportRectangle) and I could therefore plug in the rectangle data and make the viewer display a region overlay on the image. I made a test that displayed all lines at once, as you can see below:
There is quite a lot of overlap between the lines here, but this isn’t a big issue. These areas will not be displayed all at once – only one will be displayed when the user clicks on the image. If the highlighted line isn’t the one the user wants due to there being an overlap it’s very easy for the user to just click again in a slightly different location to highlight the required line.
Having managed to get the regions displayed on the image, the next step was to process click events. Despite the overlays appearing as HTML elements, each sharing the same class, it was not possible to simply use jQuery to process clicks on this class. Instead, OpenSeadragon’s clickHandler needed to be used. With this in place it was then possible to process the click, to add a new highlighting class and to grab the ID of the element, which will eventually be used to find and highlight the line in the text view.
With all of this in place I then wrote a script to export line data (IDs, coordinates) from the TEI XML and store all of this in JSON files (one per document) that can then be loaded in whenever a page is displayed. I then moved on to setting up an initial version of the website for the digital edition, setting up an initial interface, menus, pages and such things. It’s all still running on my laptop for now, but I’ve made really good progress this week and will continue with it next week.
For the remainder of the week I worked on a variety of other projects. I made some further changes to the user survey popup for the Dictionaries of the Scots Language and I added some text to Sara Pons-Sanz’s AHRC proposal document. I set up the mapping metaphor site at the new URL we will be using for the site, meaning everything is ready for the major updates that I’ll hopefully be working on in the coming weeks.
I also wrote a document outlining the ‘facts and figures’ popup for the Burns Supper map. This took some time to research, but gives a handy overview of what the popup will contain in terms of summary data and visualisations.
Finally, I returned to the Place-names of Armagh project. The existing data for this project specified location data as Eastings and Northings and I need this as latitude and longitude. When I’d previously imported the data into QGIS the markers were all in the wrong place and I didn’t know why. Thankfully Frances Kane, who is working on the project suggested that this might be because I hadn’t set the coordinate reference system to Irish Grid, and this proved to be the answer.
After making the update the data all loaded at the correct locations and with this in place I was then able to follow this answer: https://gis.stackexchange.com/a/64700 to generate latitude and longitude fields for the data and then export this from QGIS as a CSV. After that I imported the data into the CMS and all records now have latitude and longitude. There’s still more I need to do, though. At the moment the map in the ‘edit place’ page in the CMS is generated based on the supplied grid reference or ITM coordinates, which the records still don’t have, so no map displays. I think I have a script that generates this data from latitude and longitude and I’ll look into this next week. I’ll also run the data through my script that grabs the altitude for places from Google Maps by sending latitude and longitude to the service. The front-end map uses latitude and longitude directly (no messing about with grid references or ITM) so I should be all set to start deploying the front-end map soon.
Week Beginning 9th March 2026
I divided my time between four projects this week: Playbills, the Burns Supper Map, the place-names of Armagh and the Dictionaries of the Scots Language. For the Playbills project I spent some time processing the Playbill images. These were stored in hundreds of directories in many different file formats, including HEIC that is an Apple-specific format, JPEG, PNG, GIF, TIFF and even PDF. Thankfully I discovered that the command-line tool ImageMagick can process HEIC files in Windows and I wrote a script to find all image files across the hundreds of directories, move them to a single folder and convert them (where necessary) into JPEGs. It took a while for the script to execute, but it did so during Monday evening and successfully migrated 17,406 image files.
On Tuesday I then wrote a further script that picked out just the images actually referenced in our current dataset (so 1902 images) and I uploaded these to the server. I then worked on the page through which the image can be viewed. I’d already created the playbill page last week, through which the textual representation of each playbill can be accessed, and it was now a matter of adding in an OpenLayers viewer that would load in the image and enable users to zoom and pan around it. Eventually we’re hoping to connect this to a IIIF server, which will make loading the image a lot quicker, but for now the view just loads in the full, high-resolution image when the page loads. As this can take some time I added in a loading ‘spinner’. I also made the image pane as tall as the browser window and the text pane is the same height, with text scrolling if it’s longer than the pane. This ensures you can still see the image as you scroll down the text. You can also press on the icon in the top right of the image pane to view the image full screen. The following screenshot shows how things currently look, although bear in mind that the interface is still very much a work in progress – it doesn’t work very well on mobile devices yet, for example.
I also had a chat with Deven about the process for consolidating and amalgamating records using spreadsheets, and once Deven has completed this process I’ll update the data and continue to work on the site.
The Burns Supper Map took up much of the rest of the week. My work included fixing an issue with ‘guests’ not working in the short URLs, relocating a New York supper that was appearing in Barcelona to its proper location and adding ‘Source’ to the table view. I then implemented the ‘Download CSV’ option, which includes all of the fields from the table view and outputs the CSV using a purely JavaScript approach rather than connecting to the server.
I also added ‘country’ to the advanced filters, and fixed a few typos (e.g. ‘Astria’). It took several hours to implement the update, as ‘country’ is free-text and includes spaces and characters such as the colon that don’t work well in URLs. I’ve therefore had to create a new field that stores URL-friendly versions of the name and it is this that is then used for search purposes.
I then moved on to processing the supper images. This was quite nicely timed as the processing was very similar to what was required for the Playbills project (lots of different image formats with files spread across countless directories) so I could adapt the script I’d written earlier in the week. I set the script to migrate all of the images to JPEG and to copy them all to once single location with a standardised filename.
With image processing complete, I uploaded the files to the server and began work on the ‘images’ tab in the supper record pop-up, which I managed to complete. I decided to use the Bootstrap carousel to display the images, as it works very well and has support for touch controls (e.g. swiping left and right). Addin in Bootstrap (which the interface was not previously using) caused a few layout issues as some existing styles were overwritten by Bootstrap’s styles, but I was able to sort these issues out.
Where the supper has one image it just gets displayed, but where there is more than one image the carousel navigation buttons appear on the left and right of the image, allowing you to scroll through. One issue I encountered is that we have a mixture of portrait and landscape images, and this causes some layout problems. However, I’ve managed to sort this and I’ve given any extra space a black background, which I think works quite well. Another thing to note is that all of the images are loaded into the carousel at once, so for a record that has hundreds of images, or very large image files, it can take some time before everything gets loaded in. This won’t be an issue for long, though, as we’re going to pick out just a few imager per supper (some currently have as many as 200) and I’m going to resize the images (some are currently up to 16MB). Below is a screenshot of the image tab showing the image tab and accordion for one Burns Supper:
I also realised that people are going to be pretty interested in seeing the images so I added an option to the advanced filters that allows you to limit the display to records that include (or don’t include) media.
For the DSL this week I added in a user feedback popup to our test site and also at a temporary URL on the live site so that the team can test out how the popup will function. As the live and test sites use different user interface frameworks I needed to implement the popup separately for each site, but it didn’t take too long to set up.
On Friday I returned to the Place-names of Armagh resource, which I’ve not worked on for a few weeks. The PI Mícheál had gone through the lists of parishes and baronies and had noted which needed to be merged or deleted, and I spent some time processing the data, moving place-name records to different parishes or baronies as required. I also spent some time writing a script to extract the Irish forms and translations from the ‘analysis’ field (which is actually called ‘description’ in the database). The outputted table displays fields for the PID, the name and the current contents of the Description field in these columns. The ‘Irish Form’ column extracts all italic text. As there are often multiple italicised pieces of text and in such cases each is extracted and separated with a slash, and tags are displayed rather than being rendered in this and the following two columns as I’ll need the tags when I grab the data.
The ‘Translation’ column includes all of the text from the first paragraph of the description, except where this contains ‘reference’, ‘uncertain’ and ‘details’ as these text to signify that it’s not a translation. Any other text gets added to the ‘Remaining Text’ column. Unfortunately it looks like the policy of using italics for Irish forms is not used in later records, and also it looks like the HTML is malformed in a number of these later records, so all text for these rows has just been added to the ‘Remaining Text’ column. I might be able to further process things to sort some of this. But for now I’ll need to wait until Mícheál has a chance to look through the output and decides whether he wants to use it.
Week Beginning 9th February 2026
I mostly divided my time between three projects this week: Playbills, Burns Supper Map and the Place-names of Armagh. Unfortunately I was still suffering from the monstrous cold I started with last week and struggled through some of the week, but I still managed to get quite a lot done.
For the Playbills project I wrote, tested and implemented a script that extracts the data from the JSON versions of the playbill files I generated, splits this up and inserts everything into a relational database. The reason I’m doing this is to make it easier to generate canonical records for venues, plays, performers and roles, as it will be much easier to query the data and track records in a relational database.
The data I’ve extracted consists of 1902 playbill records that feature 6434 plays. These are categorised by one or more of 188 distinct genres (with ‘melodrama’ associated with 2026 plays and ‘melo-drama’ a further 10). I’ve extracted 185 distinct venues, 49498 performers, 49495 roles and 6083 contributors.
As of yet I haven’t done anything to generate canonical records, which will be the next major step, and I need to discuss things with project PI Deven before I proceed with this. For example, the role ‘Macbeth’ appears 27 times, with a further three appearances in other strings (not including ‘Lady Macbeth’) e.g. ‘Macbeth’s Last appearance’. These would need to link to one single canonical ‘Macbeth’ role. Similarly, there are 28 plays that have ‘Macbeth’ somewhere in their title, with variants such as ‘MACBETH, KING OF SCOTLAND’, ‘Macbeth; King of Scotland’, ‘MACBETH, KING OF SCOTLAND.’ In addition to ‘MACBETH’ and ‘Macbeth’ and these would need to link to a single canonical ‘Macbeth’ play.
There’s also some data cleaning that we should perform, e.g. amalgamating data that doesn’t have the same form but should be the same thing. For example, there are a lot of possible duplicates in the ‘Genre’ data. There’s ‘acrobatic’, ‘acrobatic display’, ‘acrobatic performance’ and ‘acrobatics’ all as different genres when presumably these should be the same.
I also still need to work on the performer names to split them into titles, forenames and surnames, and to ascertain gender based on titles. Venues also need some work as there are many that are the same but have slightly different text, e.g. ‘Royal Theatre, Aberdeen’, ‘Theatre Royal, Aberdeen’ and ‘Theatre Royal Aberdeen’. There’s the same issue with printers too, although perhaps this isn’t so important. E.g. ‘Keenes, Kingsmead-Street, Bath’, ‘Keenes, Bath, Kingsmead-Street’ and ‘Keenes, Bath’. It’s possible that we might be able to get some sort of AI processes to help with such tasks.
We’re also going to have to give some thought about how to handle updates to the data. I’m generating canonical records, extracting things like performer gender and generating unique identifiers for things like plays in my database, and I’ll be creating new JSON files that incorporate this new data that will then be ingested into Solr for search purposes. Therefore the data will be quite different to the original YAML files. When updates need to be made should these then be made to the original YAML files, which would necessitate much regeneration of data, or should the updates be made elsewhere, such as through the database? I don’t have an answer to this yet, but it’s something we’ll need to consider.
For the Burns Supper Map project I set up the online database for the supper data and have been working on a script that imports the data from the spreadsheets into this database. I have got everything working for the spreadsheet of the online survey, so my database currently has 308 suppers that include data for 6902 filter options.
What I haven’t been able to do yet is to import the data from the public domain spreadsheet, as this currently contains a lot of inconsistencies in how the data are recorded. The data in the filter columns (“frequency”, “category”, “toast”, “food”, “style”, “drink”, “entertainment”, “poem”, “music”, “dance”, “dress”) must exactly match the options found in the online form for my import script to work. This includes capitalisation / case and ensuring that a semi-colon is used to separate multiple items. I had a meeting with the project RA Cleo on Friday to discuss this, and she’s going to work on tidying things up.
I also wrote a script that posts the address for each record to Google Maps which then returns the latitude and longitude (something we’re going to need in order to pin the records on a map). The user inputted location data can be somewhat variable, as you might imagine, but Google Maps has generally done a very good job at identifying places from the data, and we can always tweak things once we see the locations on the map. I’m hoping to start development of the map next week.
For the Place-names of Armagh project I uploaded a large number of place-names that I’d been sent. We now have 2932 place-names in the system. I also processed the existing historical forms CSV and this has found historical forms for 1056 of these new place-names. The new place-names had additional parishes and baronies that were not already in the system and in such cases these have been created, but there are some issues, as the data appears to be somewhat messy at times and will need some cleaning. For example, there’s a ‘Forkhill’ and a ‘Forkill’ and these may be the same, there are forms with question marks and multiple forms and descriptive text, e.g. ‘Killevy/Partly in Dundonald Parish’ and ‘Armagh?/Eglish?’. These will all need separated out and fixed as required.
I also spent some time updating the CMS to convert the townland field from a textbox to a list, thus enabling multiple townlands to be associated with a place-name and ensuring each townland is only stored once in the system. This involved extracting the townlands from all of the 2932 placename records, splitting forms up that have multiple townlands in ‘x or y’ or ‘x / y’ format, storing the unique townlands and then associating the corresponding ones with each placename record. There are 997 unique townlands (although some of these may need amalgamated) and 3007 connections between townlands and placenames.
I then updated the CMS to replace the existing ‘townland’ textbox with a list of townlands as checkboxes, in the same way as parishes and baronies. We might need to rethink this, though, as scrolling through 997 townlands to find the right ones takes time. I also included an option to add a new townland when adding / editing a placename record as I’m guessing there will be more to come. This should only be used when the townland isn’t already in the list, otherwise we’ll end up with duplicates.
I also made some further updates to the CMS, namely simplifying historical forms so there is just one ‘form’ field rather than separate English and Irish boxes, and adding a flag to record whether the form is a ‘previously suggested form’ or not. I also renamed the ‘Discovery’ maps to ‘1:50,000’ as this is how the maps tend to be referred to.
Also this week I replied to a couple of emails from the DSL people about future developments and I added a new video to the Seeing Speech resource. I also investigated an issue with the Books and Borrowing website and discussed the migration of the resource to a new server with the Stirling IT people, and I generated CSV files for all of the survey answers for Speak For Yersel and sent them on to Janine Illian in Statistics, who Jennifer and I met with last week.
Week Beginning 2nd February 2026
I had a bit of a disrupted week this week, as I started feeling unwell on Tuesday morning and ended up off work sick for the rest of Tuesday and Wednesday. During this time I felt absolutely wiped out and could barely do anything other than sleep, but by Wednesday evening this had developed into a monstrous cold, the likes of which I’ve not had for several years. Thankfully once the symptoms had moved to my nose and throat my head was a bit clearer and I was able to work on Thursday and Friday, but I was still pretty far from feeling 100%.
I spent most of Monday this week preparing for, travelling to and co-presenting a talk about Speak For Yersel at the Edinburgh Futures Institute with Jennifer Smith. The talk went pretty well and it was good to meet some of our linguistics colleagues at Edinburgh, plus others involved with the EFI. I spent some of my other available time reading through and commenting on an AHRC proposal that will involve Glasgow and the Historical Thesaurus that had been sent by Sara Pons-Sanz at Cardiff University, and looking through some further place-name data I’d been sent for the Place-names of Armagh project.
Despite being off work sick on Wednesday I still managed to attend an online meeting for the Burns Supper Map project to discuss the specification document I’d prepared for the project. This was all very positive and there weren’t any major issues that anyone had spotted whilst reading through it.
For the remainder of the week I spent a bit of time investigating some issues that had been encountered when publishing pure xref entries through the Anglo-Norman Dictionary’s management system. Certain cross references were not appearing in the published entries despite being in the XML and a bit of investigation uncovered why. The entries contained cross references to entries that don’t actually exist in the dictionary. For example, Mars_2 references ‘march’, which is not an entry and respundre_2 references ‘repundre’ which is also not an entry (they both need homonym numbers added). When xref entries are published the cross references are extracted and stored, and at this point the system checks that the references are valid, and only links to entries that are valid. It is these that are displayed in the front-end, so even though invalid xrefs may exist in the XML they don’t get displayed. The ‘preview’ generates its view directly from the XML without checking validity, which is why this view doesn’t match the front-end. I ran a check and it turns out that there are around 500 xref entries that include a reference to an entry that doesn’t exist, and I passed these onto the editor who will get these sorted.
On Friday I met with Jennifer and Janine Illian, who is the current Head of Statistics, to discuss the Speak For Yersel data and what kind of additional statistical analysis might be possible. Janine is particularly interested in spatial modelling and has a keen interest in linguistics and it was really great to hear her thoughts about the Speak For Yersel data. I’m going to send her the data for all survey responses next week so she can experiment with it, and we’ve arranged to meet again later this month.
I spent the rest of my available time this week working on Deven Parker’s Playbills project, working with the YAML files, figuring out how these might be imported into Solr and how we can extract canonical records for things like venues from them. It turns out that Solr can’t index YAML files (at least not without creating a custom data importer), which is a bit of a surprise. This isn’t a major issue, though, as I can convert them to JSON, although this also proved to be trickier than I’d anticipated. Normally I’d use PHP to process data, but PHP also can’t read YAML files, at least not without installing extensions and this process seemed far too convoluted to bother with. Instead I used Python to convert the files, but this involved a bit of trial and error as I’m not used to Python and it’s bizarre insistence on whitespace being important, and the fact that if you mix up spaces and tabs to create this whitespace the scripts fall over. I got there in the end, though.
The bigger issue I encountered was with the unit of data that gets indexed. I’d previously said that we’d index entire playbill files and use ‘playbill’ as the smallest item that gets returned in the search results, but it turns out there are some problems with this, and I think indexing individual plays is going to work better. I’m still experimenting with the data and Solr’s capabilities, but initial impressions are that it isn’t very good when working with subsets of data within individual files, or more complex queries. For example, you can search the playbills for the title ‘Macbeth’ and find matching playbills. But if you combine this with another field that exists in another play in the playbill (e.g. role ‘Jacques Strop’) the playbill record will still be returned. So even though the role mentioned actually belongs to a different play in the playbill, because both pieces of information exist somewhere in the playbill it gets returned.
With my initial experiments Solr also flattened out the data – all performer names appear in one list per playbill, not separate lists per play, and it’s the same with roles. Other than the order of the items in the lists, there is nothing to connect the two. The following screenshot shows one playbill record indexed within Solr (just using Solr’s default post and without customising a schema). You can maybe see how Solr has flattened things out, resulting in data being lost (e.g. which performer belongs to which play).
I then tried to index the data at play level, adding in a play ID and also any playbill level data (thus ensuring it’s still possible to search for date, venue etc). You can see the results in the following screenshot, which includes 5 separate records.
Here at least it’s possible to tell which performer / role belongs to which play. But performers / roles are still only connected by their position in the lists. Record 5 is a duplicate I made of record 4, but I deleted the ‘role’ text for one performer to see what would happen. And Solr indexed the record as it was, with 5 performers and 4 roles, so based on list order ‘Miss Newton’ is now ‘Landlord’ and not ‘Marie’, and ‘Mr. Watkins’ now had no role.
After further investigation I realised that it is possible to get sole to properly index nested data (see https://solr.apache.org/guide/solr/latest/indexing-guide/indexing-nested-documents.html) although instructions on how to actually import nested data into Solr are pretty thin on the ground – you can’t just use the default ‘post’ command as this flattens all data. I ended up following another tutorial (see https://docs.arenadata.io/en/ADH/current/how-to/solr/solr-index-nested-docs.html) and importing the data using the Solr admin interface. This thankfully worked, as the following screenshot demonstrates. You can see that individual performers are directly associated with roles.
There’s still a massive amount to do with the data, though. I need to extract unique venues, plays, performers and roles and assign IDs to them to enable them to be searches for. I decided that it would be easier to manage such processes via a relational database, so on Friday and mapped out a structure for the playbill data and bean working on an import script that would process the JSON files. Lots more to do in the coming weeks!












