Category: Bids
Week Beginning 25th August 2026
This week I continued with the interactive map for the HiMuJe Malabar project and made some good progress with its development. All subtypes now have an associated icon that is used on the map when the ‘type/subtype’ categorisation is selected, and I also reworked the colours of these, as you can see in the following screenshot:
After last week’s frustrations with the map legend, I reworked it to divide it a little better into sections by type, which you can also see in the above screenshot. When a location popup is opened the corresponding subtype icon also now appears in the popup header. I also added an option to the ‘Home’ menu to make all marker labels permanently visible on the map (until you turn them off again). With all locations on this does clutter up the map, but I think it’s a useful option for people who are looking for a place but aren’t entirely sure where it is. With labels on you can see the names of all places without having to hover over each marker in turn. This works better when the map is filtered in the legend. I also intentionally set the map so that I you set the labels to on and open the travel route menu the labels are removed as the travel routes would be very difficult to see with the labels on.
I also updated the colours for the other categorisation types to hopefully make them work a little better. When the ‘number of references’ categorisation is selected the marker colours are now a gradient. It’s also now possible to cite a specific stage in a travel route. This appears in the citation text, e.g. “HiMuJe Malabar Digital Map: Watercolor map, travel routes menu, showing the travel route ‘Brahmins’, with unrelated places hidden, viewing stage 2 of the route, placenames categorised by language. 2026. In HiMuJe Malabar. Glasgow: University of Glasgow. Retrieved 28 August 2026 and following URL opens the route at Stage 2. This option will mean I’ll be able to update the location popup to add a new tab that lists all of the stages in routes that the place appears in, and give links to each of these. I haven’t done this yet, though.
Also this week I spent quite some time working for the Dictionaries of the Scots Language. This included attending an online meeting to discuss the structure of the new site, at which some fairly major revisions to the new design were proposed. I also engaged in a couple of email conversations about the updates that still need to be undertaken before the hopefully launch the new site later this year. In preparation for all of this work I also went back through all of my emails from the DSL team and my notes to create a bit ‘to do’ list containing everything that needs done. It’s going to be a busy couple of months.
Also this week I met with Craig Lamont and his RA Emily Hay to discuss the Life Writing project. They’re going to be compiling a bibliography of Scottish life writing and I’m going to develop the systems for this. I did look into using existing bibliography software but nothing I looked at did exactly what they are after so I’m going to create a simple CMS for them to use myself, with a front-end consisting of search and browse options and a map interface sometime next year. I spent a bit of time after the meeting going through the sample spreadsheet they had prepared and sending a list of questions about the data.
I also met with Deven Parker this week to discuss an AHRC proposal she’s involved with. I can’t say too much about this, but it’s related to the work we’re doing on the playbills and I would be involved in some technical capacity. I also had a chat with Garrick Allen about a proposal he’s trying to get funding for after an earlier AHRC submission was unsuccessful. I read through the documentation and gave feedback, and I’ve also put Garrick in touch with my colleague Luca Guariento who is probably going to provide the technical assistance.
Other tasks I carried out this week included investigating an issue with the Anglo-Norman Dictionary (which turned out to be an issue with the data and not the system), making some final tweaks to the Burns Supper Map before the project RA finished working on the project, and generating some Books and Borrowing data for a publication Katie Halsey is working on.
Week Beginning 17th August 2026
I continued to work on the DOST Auld Laws project on Monday and Tuesday this week. My first task was to implement the tag selection for the Advanced Search. As discussed last week, this section allows users to either include or exclude tags from their searches, or limit their search to only look at the contents of one of more specified tags. Luca has already created the XML search options that would allow this, and my job was therefore to process the user’s selections, format the query and connect to the XML. I managed to complete this on Monday and it’s now possible to construct a query such as “find all occurrences of words beginning ‘ȝe’ in ‘Edzell Doouments’ excluding any text found in ‘Expanded Forms’, Aitken Notes and Other Notes”. It’s also possible to press on the ‘Refine your search’ button to return to the search form and the search criteria will be remembered in the form. I also added in a ‘Clear’ button.
I continued working on the site on Tuesday. I added a ‘Cite’ option to the search result page, which allows users to share or cite a particular page of a search, including the ordering. I also added a help popup about the tag selection to the search form and I integrated the ‘show all Aitken notes’ query with the Advanced Search. This is a special case that returns the full contents of all notes by Aitken across all or a selection of documents without the need to supply any search text. Information about how to do this appears in the help popup as follows:
“To retrieve all Aitken Notes enter a dash (-) into the search box, select documents (if required), set Aitken Notes to Limit to content in this tag and press the search button. This will display the full contents of all Aitken notes in your chosen documents.”
The Aitken Note results are a bit different to the regular results as they return the full contents of each tag, not a KWIC. Therefore all of the text is highlighted in yellow and the sorting by left and right of the term doesn’t do anything, as there is nothing left or right.
Whilst working on the Advanced Search I’d spotted that the advanced search KWIC (KeyWord In Context) was not crossing line boundaries, meaning that when a work is found at the beginning of a line no contextual text is returned before this. This is different to the Quick Search, which ignores line boundaries and is not ideal, as it means the search results are inconsistent and the quick search actually gives better results than the advanced search. Thankfully Luca was able to find a solution to this and to make the advanced search results consistent with the quick search.
I’ve now pretty much done all I can do for this project until I hear back from Joanna, who is going to supply ancillary content and edited XML files.
I spent a lot of the rest of the week working on the HiMuJe Malabar project, continuing to develop the map. On Thursday I attended a meeting with Ophira, Renu and project Co-I Ines, who is in Glasgow from Germany for a while to work on the project. We had a very productive meeting and we now have a much clearer idea about how the travel routes and other aspects of the map will function.
In terms of actual development of the map, this week I fixed an issue with the travel routes whereby when multiple travel routes are selected, deselecting one would remove all locations, even those that were associated with another active route. This took quite some time to sort out, unfortunately, but it’s fixed now and the routes work a lot better.
I then set about making it possible to share or cite the selected travel routes, also noting whether unrelated places have been hidden or not. This required some major reworking of the code to ensure that these options are added to the URL and when the page loads that the options are taken from the URL and processed. The citation text also needed to include information about the selections too, and any location selected when the route is active can also be cited, e.g.
“HiMuJe Malabar Digital Map: Satellite map, travel routes menu, showing the travel route ‘The migration of Jews and Christians in the Qissa’, with unrelated places hidden, placenames categorised by type / subtype, placename record for Mecca. 2026. In HiMuJe Malabar. Glasgow: University of Glasgow. Retrieved 19 August 2026”
At the meeting we decided what icons to use for the location categorisation so on Friday I returned to the map legend to rework this. I started off by slightly reworking the data to ensure that all locations had both a type and a subtype, as a few locations only had the former. This meant creating a new ‘Area’ type and making ‘Region’ a subtype of this, whereas previously ‘Region’ was the type. The reason for doing this is so that the legend could be grouped by type, giving each type a count of the total number of locations contained across all subtypes and to include a checkbox for each type that when pressed on would select and deselect all subtypes.
That was the plan, anyway, but unfortunately I had a very frustrating day and didn’t manage to get this working by the end of it. The difficulty is that Leaflet generates and processes the legend internally based on the map layers and trying to hook into this and change the default behaviour is very difficult. I managed to get the ‘Type’ checkboxes to select and deselect all corresponding ‘Subtype’ checkboxes but doing so was not actually triggering any changes on the map – the corresponding map layers were not being affected when the checkboxes changed even though manually pressing on the checkboxes did trigger the layer change. I’m afraid I ran out of time with this and didn’t manage to find a solution. For now the types are just headings in the legend without associated checkboxes, which is a bit of a shame but I just don’t have the time at the moment to look into this further.
Also this week I made a few more tweaks to the data for the Burns Supper Map, including adding in a new batch of images, discussed migrating a couple of sites to our third-party hosting supplier, read a document ahead of next week’s DSL meeting and read some information Deven Parker had sent me about a funding bid.
Week Beginning 10th August 2026
I was back in Glasgow and back to a five-day working week this week, as the summer holiday period drew to a close. I spent a fair amount of time this week continuing to work on the DOST Auld Laws project. Luca has been working on the XML queries I’d specified for the advanced search and I was able to test them out and give feedback on them. These are mostly all working as I’d hoped, which is really great, and I was able to begin working on the front-end aspects of the advanced search, with the aim of connecting all of this into the queries Luca had created.
However, I also had to spend quite a bit of time working with the source files, as the project PI Joanna Kopaczyk-McPherson had realised that two of the eight documents should really be split up into smaller sections. This was no straightforward task, as not only did it require the XML documents to be split into smaller sections, but many other aspects needed to be updated. The document IDs needed to be changed, which mean IDs used for image filenames throughout the documents also needed to be changed, with the filenames of the actual images also then needing to be updated too. For page navigation in the site data is stored in a database and this also needed to be updated, and the XML files stored in eXist for search purposes also needed to be updated. There were twelve steps I needed to follow for each required split, which took some time, but thankfully the process went pretty smoothly and we ended up with 13 documents in the site instead of the original 8.
There were further issues to come, as Joanna has now noticed that the documents need further edits and tweaks, and not just minor changes to text but structural issues such as the insertion of omitted lines. This is going to be very difficult to do as lines are linked to coordinates in the images, which was all handled via the Transkribus tool. We exported the XML files from Transkribus month ago and many major changes have been made to the files since the export so it’s not going to be possible to re-import them. Manually creating new lines in the XML files as they are now will not have the connections through to coordinates in the corresponding image files, so we’ll end up with inconsistent data. It’s not a great situation to be in, and ideally all editing of the documents should have been completed in Transkribus before we exported the files, something I’d mentioned back when I undertook the export process months ago. We haven’t reached a decision on how best to handle this yet and we’ll continue to discuss the options next week.
Despite all of this I did manage to work on the advanced search, creating the advanced search form which, as specified, features a textbox where you can enter some text, a list of documents from which you have the option of selecting the ones you’re interested in and a list of tags that you can either include, exclude or limit your search to, as you can see in the following screenshot:
As of yet I have not implemented the tag limits, but the limit by documents is operational. This connects through to Luca’s new eXist-db queries to perform a search limited by documents and as with the quick search, you can also sort the results by words to the left and right of the term, although I’ll need to get Luca to look into how the KWIC is generated for the advanced search results as they don’t seem to cross line boundaries, unlike the quick search. The advanced search results display information about the documents you’ve selected and if you press on a search result to load the corresponding page the link back to the search results takes you to the right place. There’s still lots to do. The limit by tags is the biggest thing and will probably take quite some time. I also still need to add in an option to cite a specific search result page and add in an option to refine your search, which will remember the options you previously filled in when you return to the search form. An option to clear the search also needs to be added. I’ll continue with this next week.
Also this week I participated in two meetings about the place-names AHRC proposal I’m involved with, and this is coming together very well. We now have an outline proposal completed and pretty much ready for submission.
I also spent a bit of time continuing with the travel routes for the HiMuJe Malabar project, adding in a few more travel routes that had been prepared and made a few tweaks to the XSLT that generates the entry HTML for the new DSL interface, fixing some issues with the layout of the new ‘combs’ sections.
Week Beginning 3rd August 2026
I worked a total of four days over the past two weeks, and was on holiday for the remainder. During this time I had a meeting to further discuss a place-names related AHRC proposal I’m involved with. I can’t really say much more about it at this stage, but the proposal is coming together. I also had to spend some time working with IT Support and Luca to figure out why our local server kept going offline repeatedly. It looks like this was caused by the server getting swamped by requests from one particular source (almost certainly bot or AI) and thankfully IT Support were able to block this, after which the server was stable again. It’s something we’re going to have to keep looking out for in future. I also spent a bit of time working with Luca to get automatic WordPress updates working on our local server, as the way sites had been set up meant that the setting was not working. Luca managed to find a solution to this, which is really great.
I also spent a bit more time working on the Burns Supper Map, creating a record for it on this site (see https://digital-humanities.glasgow.ac.uk/project/?id=156), adding more suppers that had been submitted via the survey and making some requested edits to existing suppers. I also set up access to Google Analytics for the other two members of the project team.
In addition, I investigated an issue with the Scots Syntax Atlas after someone suggested that the linguists’ atlas was looking somewhat blurry. I managed to figure out why this might be the case, although I’m not entirely sure whether this is a new issue or if the markers always looked like that. I contacted the project PI and suggested a couple of updates, but I haven’t heard back yet so will need to wait and see what she says.
I spent most of the remainder of my time working on updates to our new test interface for Dictionaries of the Scots Language and working through the list of outstanding items for the DOST Auld Laws project. For the DSL I completed the updates to the bibliography page that I began working on a couple of weeks ago. I implemented pagination of the entries associated with bibliographical items, with navigation bars appearing above and below the entries, with 20 appearing per page and ‘jump to page’ buttons also appearing, just like with the search results. This works pretty well, but is somewhat cumbersome for someone like Sir Walter Scott, who is referenced in 2861 entries, split over 144 pages. We don’t have this issue with the search results are these are capped at 500 (25 pages) so we might need to think of other ways of handling this.
I also ensured that headword searches that don’t yield any results automatically perform a fulltext search for the term supplied. This works on the live site, but only when both dictionaries are selected. With the new site there are separate quick searches for SND and DOST so the additional search wasn’t being triggered. It is now, as is the advanced headword search for both dictionaries.
I also made tweaks to the DSL’s new regional map based on feedback I’d received – adding in some content where we previously had placeholder text and ensuring the ‘About’ popup didn’t disappear off the bottom of smaller screens and a few other small updates. I then began to look at the ancillary pages and how we can make them look a bit nicer. I spent a bit of time on the ‘Word of the Week’ page and liaised with William Ashford, who is responsible for such content about this, and further updates that were going to make to the ancillary content closer to the launch date of the new site (which will hopefully be in November).
For the DOST Auld Laws project I added the top navigation bar that will link the site in with the SCOTS Corpus and CMSW. I added in copyright information and added facilities to download page images and the XML files for each document. These being up a pop-up asking for users to abide by the license before leading to the actual content, which hopefully won’t be too annoying. I also added in the ‘cite’ popup to all document pages, which took a little time to implement, and added in Google Analytics. I also made the image thumbnails on the document overview pages smaller and placed them in a collapsible section that is closed by default, plus I removed the introduction to the documents page, as this will be covered by the homepage.
I also removed some pages that didn’t have content (e.g. blank pages) from the beginning and end of some of the documents and I added a feature to turn off and on the line highlighting feature. The highlighting feature allows the user to click on a line of text in the image or text for a page and for that line to be highlighted in both the image and the text, which is pretty nice. Unfortunately the line highlighting gets in the way of the image viewer’s zoom and pan functionality on touchscreens, making it a somewhat unreliable and frustrating experience. This new feature removes the option to ‘click’ on a line, meaning pointer events are not intercepted and make their way reliably through to the image viewer, which works much more smoothly.
Also this week I had an email conversation about user feedback and walkthough videos for the STAR resources and booked my accommodation for the DHC conference in Sheffield. Next week I’m back in Glasgow and back working a full week, with summer holidays all over.
Week Beginning 13th July 2026
This was a four-day week for me as I’d taken Friday off (and I will also be off for the first three days of next week). I finally managed to assign some time this week to implementing updates to the new Dictionaries of the Scots Language interface based on feedback that had been sent to me earlier this year. I spent most of Monday and Tuesday working on this.
I can’t share any screenshots of the new interface at this stage, but I updated the search results box in the entry page to remove the tab for the other dictionary when performing a search (quick or advanced) for a specific dictionary. This avoids misleading people as it otherwise the tab displays zero results for the other dictionary when in fact it just means the dictionary wasn’t actually searched. Now when a quick search is performed, the other tab heading is replaced by a link to search the other dictionary. Pressing on this performs whatever search you’ve executed (quick or advanced) on the other dictionary.
I also tweaked the font colour of the inactive tabs. I realised that the white text on grey made it look like the tabs were disabled, when they’re not, they’re just inactive. I therefore made the font darker, which I think works a lot better. I then added in a button that scrolls the page to the search / browse box. This appears above the entry header (and in the ‘sticky’ header that appears as you scroll the page) and only appears on narrower screens (where the infobox appears below rather than beside the entry text). Where the entry is in the search results the text is ‘Scroll to results list’ with a down arrow. Otherwise the text is ‘Scroll to browse list’. The DSL team had requested that results term highlighting should be off by default, so I made this change too.
I then began to rework the bibliography page based on feedback. We’ve decided to go with the version of the bibliography page that displays the quotations from any associated entries in addition to the headwords and links through to the entry pages. This required some reworking of the API so that rather than returning individual citations, it brings back entries with each associated citation as part of this. This means that multiple citations for an entry no longer appear as separate items in the list but are grouped by their entry, much like the quotation search results. I also updated the count above the citations to display both the number of entries the item appears in as well as the number of citations (e.g. ‘Cited 718 times in 597 entries’) and each citation also includes its date as well now. However, the order of citations needs to be the order they appear in the entry and not date order, otherwise the links through from the citation may end up taking you to the wrong one in the entry (as I discovered when I set things to date order).
The display of entries and citations is not exactly identical to the search results. There is no sparkline as unfortunately this is not included as part of the bibliography data and I’d need to rework the database and API in order to include them. Also, the quotations are in a larger font than the ones in the quotations search results as they seemed a bit small, and the headword isn’t highlighted in the quotations as this is something performed by the Solr search engine and isn’t available as things currently stand for the bibliographies. I still need to add in pagination, which I didn’t have time to work on this week but will hopefully implement soon.
I also participated in a Teams meeting with the DSL this week that involved their interns reporting back about user engagement and observations about the website. It was very interesting to hear their feedback and it will give us lots to think about as we continue to improve the resource.
Other than working for the DSL, I also made some last-minute updates to the Burns Supper Map before its official launch over the weekend. The resource is now available for anyone to use at https://burns-supper-map.gla.ac.uk. I wasn’t able to attend the launch as I was away on holiday but from what I’ve heard it was a great success.
I also spent a bit of time fixing an issue with the Anglo-Norman Dictionary where certain entries that had an apostrophe in their headwords (e.g. j’) were not loading while others were. The reason for the discrepancy was because some headwords had curly apostrophes and other had straight ones. The straight ones were getting encoded (e.g. “j'”) and were then not getting found. Once I managed to figure this out I was able to fix the issue.
On Wednesday I had a lengthy online meeting for the HiMuJe Malabar project to discuss the interactive travel routes. The meeting lasted about two and a half hours but it was worthwhile as we all have a much clearer idea of how to proceed with the routes now. I just need to wait until the team sends me some initial travel routes using the spreadsheet template I sent them and then I’ll be able to continue my work on this.
On Thursday I participated in a Teams call with colleagues from Nottingham and Cardiff Universities to discuss a place-names proposal that I am likely to be involved with. I can’t really say much more about it at the moment, but it’s all sounding very interesting. I spent a few hours after the meeting writing a document containing some initial thoughts about how the technical infrastructure for the project could work.
I was then off on holiday on Friday and I won’t be back at work again until next Thursday.
Week Beginning 18th May 2026
This was another week of many projects, the first being the Burns Supper project. This week I finished implementing all of the data visualisations on the ‘facts and figures’ pop-up, with a mixture of bar, column and pie charts depending on how many data types there are. All charts are ordered by number of suppers, other than the pie charts. I also added in a statement above each visualisation about the number of suppers that supplied the data and a note when a supper may have more than one type of data. The screenshot below shows part of the popup with three visualisation types (mostly) visible:
I haven’t updated the narrative text to update the percentages to be based around the number of suppers that supplied the data rather than the total number of suppers and to include additional explanatory text, but I’ll do this next week.
For the HiMuJe Malabar project I have now incorporated the updated JSON file that contains cats and subcats into the map. I have also incorporated the GeoJSON polygons that project RA Renu has created so far. Where a place has a polygon this is used instead of a marker, and polygons feature a tooltip on hover-over in the same way as markers. I added the type and subtype to the tooltip as well. I also made a start on the categorisation of places based on cat and subcat, although this still needs some work. I added in a ‘legend’ box on the right of the map and have split the places into layers based on their cat and subcat. Each of these appears in the legend, and you can turn each layer on or off. You can also select or deselect all layers, which I find quite useful (e.g. deselect all then only add in the layers you’re interested in). There is also a count of the number of places in each layer as part of the layer label. Note that the cats and subcats only appear if there is at least one place that has latitude and longitude assigned to the subcat, which is why (for example) ‘Mosque’ is not currently listed. Below is a screenshot showing the categorisation options that are currently in place:
There’s still quite a lot to do and some issues that need to be addressed. At the moment only places that have a subcat appear. I still need to implement the top-level cats, so for example ‘Sri Lanka’ that has cat ‘Region’ and no subcat is not currently appearing on the map. I’m hoping to update the legend to give it a two-level hierarchy (as mentioned in my specification document) but I didn’t have time to implement this. Also, marker and layer colours are currently arbitrarily assigned. I think we will eventually have icons as well as colours, but in addition to this I will update the colours, probably to have different shades for each subcat in a cat. I also need to ensure that the polygons always appear behind the markers as this is not currently the case. E.g. if you turn the ‘Region: Kingdom’ layer off and then on again the layer is then added to the front, meaning it’s no longer possible to press on any markers located within the polygons. Also, there will be several other categorisation options in addition to cat and subcat (as discussed in the specification document) and these still need to be implemented.
Also this week, I returned to working on the Playbills project for the first time since before Easter. This included writing an executing a script to merge genre classifications, which reduced the number of genres from 188 to 69 and reassigned any plays that were assigned to a deleted genre. I also executed my scripts that process performers to extract individual names and titles from the free-text ‘full name’ column and assigned gender based on the title. My scripts process individual performers and also split multiple performers into individual records, resulting in the number of performers going up from 49498 to 52146. In these cases the newly extracted performers have also been assigned to the same role as the existing row. Of the 52146 performers we now have a Male or Female gender assigned to 51724, which I think is pretty good. I also updated the playbill and play pages so that the extracted names and genders now appear, and multiple performers for a role all now appear. I still need to deal with the ‘unprocessed’ performer spreadsheet, which is probably going to need some manual work as this contains the ‘edge cases’ that my processing scripts were unable to tackle. I hope be able to continue with this next week.
My fourth project of the week was the DOST Auld Laws project, for which I continued to work on the search facilities. This has involved setting up an Exist-db XML database and learning how to run full-text queries on the documents contained in it and how to generate KWIC (keyword in context) snippets for the results. There’s still a lot to do, but an initial version of the quick search is now operational, as the screenshot below demonstrates:
You can also use an asterisk wildcard to represent any number of characters, for example ‘wyn*g’ finds all terms beginning ‘wyn’ and ending in ‘g’. A question mark wildcard can also be used to represent a single character, e.g. ‘r?cht’ matches ‘richt’ and ‘rycht’. I haven’t added in pagination of results yet, so they currently all appear on one page. Each result features the document name and the page number where the result is found, plus the snippet with the term highlighted. Pressing on the snippet loads the page with the line where the term is found highlighted in both the image and the text.
It’s definitely something of a milestone to get the quick search working, but there is still a lot to be done. The highlighting of the term in the snippet is not currently working properly when the term includes characters in <expan>. The search finds the term, but the highlighting is only getting applied to the part of the term before the <expan> tag. Also, Exist-db uses the Lucene full-text engine and this does not support wildcards at the beginning of terms, so while a search for ‘grant*’ finds all terms beginning with ‘grant’, it’s not possible to find all words ending in ‘*ting’, for example. I’m not sure how big an issue this is, but if such functionality is required I’ll have to investigate an alternative. Also, phrase searching is not currently operational. I have a query that should be able to handle this, but I haven’t had time to add this in yet. It’s been quite tricky to implement as phase searches need to cross line and page boundaries whilst still returning the line and page IDs where the start of the phrase is found.
In addition, Boolean searches are not yet working (e.g. term1 AND term2, or term1 OR term2). However, we might want to consider how useful these would be and how they would work. Currently the search looks for occurrences of the term in entire documents and returns snippets showing the context of the term. If a Boolean search is using an entire document as its source then how useful would it be if (for example) term1 is found on page one and term2 is found on page 29? I’m just not sure how helpful this would be. I guess an OR search would be useful in finding variant spellings?
I also still need to add in results traversal to the document page, so you can navigate directly through results when looking at one specific result, and I need to add in a link back to the results page from the document page. There’s also still a lot of work to do on the user interface, which is still not finalised and I’ll work on this (e.g. colours, fonts, layout) once I’ve finished with the search facilities.
My fifth and final project of the week was the redevelopment of Mapping Metaphor. I didn’t have much time left to spend on this, but I did manage to tidy up the loose ends from last week. Where a combined metaphor has both OE data and E data and the E data does not begin in the OE period the timeline in the visualisation card popup now includes OE selected and highlights the period when the E data begins. This second highlight currently uses the green used for the visualisation background but can be changed. We should also maybe include some explanatory text so users understand why two periods are selected. The following screenshot shows the timeline with two periods highlighted:
I also updated the card popups so that both E and OE IDs appear in the header of the card (where applicable) now, as you can also see in the above screenshot. Also, where a metaphor connection exists in E but not OE (or vice-versa) counts of lexemes in each joining category in the other period are now displayed.
Also this week I had an email conversation with Craig Lamont about an online resource for his new research project and arranged for a subdomain to be set up for it. I also updated the DSL survey so that it only appears on mobile devices and responded to a request from Kirsteen McCue regarding stats for song downloads on the Editing Burns and Burns Choral websites.
Week Beginning 4th May 2026
This was a four-day week as Monday was a bank holiday. I spent most of Tuesday and Wednesday continuing to work on the DOST Auld Laws project, working with the XML files. The XML files generated by Transkribus and exported by the tool as TEI XML contained many elements that were not valid TEI elements, such as the custom <Aitken> element that had been applied to notes added by A J Aitken. My first task of the week was to write and apply transformations to the XML files to convert them into fully valid TEI. I achieved this using XSLT, which is a language I find very unintuitive and frustrating to work with, no doubt exacerbated by the fact that I don’t work with it very often. I struggled to get any transformations to run initially, and ended up turning to AI to figure out why my scripts were not working. In this instance AI proved to be extremely useful as it identified what the problems were (e.g. I hadn’t declared the correct namespace or used it when writing the rules) and really helped me to understand how everything fitted together. I still wrote the bulk of the code myself, but AI was very helpful in identifying errors or issues. By the end of Tuesday I had generated (and checked) a collection of TEI files that successfully validated in Oxygen, which was a good milestone to reach.
On Wednesday I then worked on the front-end for the project, figuring out how to transform the valid TEI XML into HTML for display in the ‘text and image’ and ‘text only’ views of document pages. As this was once more using XSLT I enlisted the help of AI to figure out specific issues that I was unfamiliar with. The biggest of these was how to pick out and process one single page from a document’s XML file based on the <pb/> element. I had no idea how to achieve this, and despite this being a fairly fundamental issue when processing TEI documents I didn’t manage to find any useful information online. However, AI came up with a solution (and equally importantly an explanation) in a few seconds and I was then able to incorporate this into my code, transforming the contents of one page, whose ID was passed to the XSLT script as a parameter, to HTML with a variety of styles applied to the various elements.
I also updated the ‘click on a line in the image’ feature, and it’s now possible to deselect the line if it’s already selected – previously once you’d highlighted a line you couldn’t get rid of the highlighting, only move it to a different line but now if you press on the highlight it’s removed. The lines are also now connected to the text view – pressing on a line in the image also highlights the corresponding line in the text. You can also press on a line in the text to highlight it and the corresponding line in the image.
I also updated the height of the text pane so that it matches the height of the image pane and if the text is longer the pane scrolls. This ensures that if the text is very long it’s still possible to see the image, rather than having the entire page scrolling, which may result in the image not being visible when you’re at the end of the text. It is how we did things in Books and Borrowing, but having a scrollbar in a section of the page in addition to the browser’s scrollbar can be annoying for some people so I might revert to the previous layout depending on feedback. Here’s a screenshot showing the transformed text, a highlighted line and some of the formatting that’s been added:
Also this week I continued to work on the Place-names of Armagh project. I wrote and executed a script that generated Irish grid references and ITM values for all places based on their latitude and longitude, a task that I completed using the ‘GridRefUtils’ scripts as detailed here: https://www.howtocreate.co.uk/php/gridrefapi.php. I’ve used these scripts before on previous place-name projects and they’ve been hugely useful. I then ran a further script to generated altitude for the place-names by connecting to the Google Maps API, so we now have complete geospatial data for all of the places (other than the 386 that didn’t include Easting and Northing data in the original spreadsheet). I also added in a new parish and barony in a different county that one place-name requires, and had discussions with the team about the splitting of analysis data into Irish forms and translations. Removing the Irish forms from the translation field is going to be rather tricky to automate as the Irish forms often form part of the translation. It’s looking like I’ll need to automate the transformation of some of these, with the rest then requiring manual intervention.
I also continued working on the redevelopment of the Mapping Metaphor resource to create a combined English and Old English map, something I began last week. This week I managed to get the combined view of the drilldown of the visualisation working. The counts represented by the yellow circles also use the combined data, and these are also displayed in the pop-up card, for example, if you select 1K02 Creation and press on the yellow line or circle for 1B Life. In the combined map there are connections to 6 categories in Life whereas there are 5 in the E map and 2 in the OE map (one of the OE ones is also present in the E map, which is why the combined total is 6). The combined pop-up card also now lists the number of OE lexemes in addition to the E lexemes, for example “1K02 Creation 1088 lexemes / 127 OE lexemes”.
I still need to implement the combined visualisation view of the search results, and also the card pop-up between individual categories, which is going to take quite some reworking (possible direction changes, combined example lexemes, timeline updates etc). I’ll hopefully find some time to continue with this next week.
Also this week I processed and added some more images to a few Burns Suppers, had a chat with Garrick Allen about a research project he’s wanting me to be involved with later in the year, and contacted Lindsay Balfour about a proposal she’s writing that will have some technical requirements. I also made a couple of tweaks to the Thesaurus of Old English website after Fraser go in touch with some suggestions, made some tweaks to the survey popup on the Dictionaries of the Scots Language website, gave some mapping advice to Renu of the HiMuJe Malabar project, and created an alternative version of one of the Anglo-Norman Dictionary’s textbase documents that strips out all Latin text.
Week Beginning 27th April 2026
I had several meetings this week, the first being to discuss the development of the map for Ophira Gamliel’s HiMuJe Malabar project. In the run-up to the meeting it became clear that no-one had read the specification document I’d written regarding the development of the map, which I’d sent out several weeks ago, but thankfully after I raised this it was distributed before the meeting and we had a very useful discussion about the map and how to prepare the data for the map. It feels like some real progress has been made in our joint understanding of the data and what needs to be presented, and following the meeting I sent out a proposed structure for the JSON data that will be generated from the digital edition files and a proposal for how we would document travel routes on the map. I also met with project RA Renu a couple of days later to discuss geoJSON data and the creation of polygons for certain locations in the data.
My second meeting was with the Place-names of Armagh people, and this was also very useful. We discussed some additional explanatory text found in the original data that I hadn’t spotted before, the splitting up of the existing ‘analysis’ data into Irish forms and translations and the creation of the front-end. Now all I need to do is actually find the time to work on all of this. My third meeting was with the Burns Supper team, at which we went over the ‘facts and figures’ popup that I will be developing. We also discussed the launch of the resource and a demo that was given on Monday and was very well received.
I spent a fair amount of time this week working on the DOST Auld Laws project, beginning with an analysis of the XML files in order to work out what non-standard elements generated by Transkribus needed to be converted into which TEI elements. It took some time to go through everything but it was very useful to get it all documented and subsequent email discussions with Joanna and Pia were really helpful.
I then ran a script to update the XML files to change the filenames, add in IDs and replace the referenced image filenames (both in the XML files and the image files themselves) with a more rational naming structure as I’d previously discussed with the team. For example, the file with the title ‘115 – Peebles B. Rec.’ was given the ID ‘peebles-b-rec’ and image files and references were renamed from (for example) ‘0001_115 001.jpg’ to ‘peebles-b-rec_001.jpg’. With this in place I then arranged with Luca for the images to be uploaded to the IIIF server, after which I could begin working on the front-end on the server, as opposed to on my laptop as I’d been doing previously.
As part of this work I also decided to migrate the document data I’d extracted from the XML from JSON to a MySQL database. The reason for this was that some of the document JSON files were rather large (a couple more than 6MB) and it seemed rather inefficient to load all of this data in just to work out counts of the number of pages and such things. With all of the data in a relational database it’s much easier to query and return just the data that is needed.
During the week I completed work on an initial version of the image parts of the site. The user interface is by no means complete and is fairly rudimentary – we will decide thinks like fonts, colour schemes, illustrative images and ancillary content later on. For now there is just placeholder text on all but the ‘Documents’ pages and the search option doesn’t work yet. However, it is possible to browse the documents and all facsimile images contained in each. Pressing on the ‘Documents’ menu item brings up a sub-menu listing the documents, plus an ‘overview’ page which is currently empty. If you select a document you’re presented with an overview page. This can feature a description of the document (there’s just placeholder text for now) a link to open the document at the first page and a randomly selected image displayed to the right of the description for illustrative purposes.
Beneath this is a section through which you can see a count of the number of pages and access a list of thumbnails of every page. For some of the longer documents it can take some time to load in all of these thumbnails. We can maybe have this section hidden by default, or possibly paginated. Pressing on the ‘open document at first page’, the random image or any of the image thumbnails opens the page in question, a screenshot of which is shown below:
I’ve borrowed much of the layout of this page from the work I did on the Books and Borrowing project. The page features a navigation bar at the top and bottom, through which you can navigate to the next or previous pages, or use the ‘jump to’ feature to select a specific page. There are three views of the page: an image and text view which displays both side by side and then text or image only views. When you select a view it is remembered as you use the navigation options.
As of yet there is no content in the text view (processing the XML will be my next task) so for now it’s all about the images. You can zoom and pan the image or open it full screen. You can also press on a line to highlight it. Eventually this will allow highlighting of the text from the image and vice-versa, but for now all that happens is the section of the image is highlighted. As mentioned in an earlier post, I adapted an existing ‘simplify coordinates’ script (https://github.com/dariok/page2tei/blob/master/simplify-coordinates.xsl) to generate simple rectangular sections for each Transkribus polygon, but we may need to have more complex shapes as in some pages (particularly the handwritten ones) the boxes are not very accurate. I’ve also spotted an issue on touchscreens (on my Android phone using Firefox, at least) whereby the boxes for the lines stop the ‘pinch to zoom’ feature working, meaning zoom will only work using the icons in the top left. I’ll need to investigate this further.
My next task will be to work with the XML, firstly replacing the Transkribus tags with valid TEI ones and then generating the text view for the website, which will link each line to the image. Once this is in place I’ll begin to think about the search facilities.
On Friday I spent a bit more time working on Sara Ponz-Sanz’s AHRC application and spent the rest of the day beginning work on the redevelopment of the Mapping Metaphor website. It took a while to refamiliarize myself with the code and to work out how things function, but once I made some progress with this I decided to start work on a new ‘combined’ version of the map. This will display the combined E and OE data but will not initially allow you to switch the display to just E or OE. There is a LOT of work that needs to be done to get this working fully, and I only made a start on things this week. For now only the visualisation works, and only the top level view of this works.
Working out the totals for the visualisation is not as simple as just getting the E and OE data and adding them together, as this would duplicate data for metaphors that appear in both sets. Instead, the new version of the visualisation grabs all of the E data and then only adds OE data for connections that are not already present. So for example, in the new visualisation pressing on ‘1K’ displays a connection to ‘3E’, which is not present in the E map but is present in the OE map.
In the combined map with strength set to ‘Strong’ and with ‘1K’ selected, pressing on the yellow line or circle for ‘1B’ the card shows 18 connections. If you do the same for the E map you’ll also see 18 connections, while the OE map shows 3 connections. The reason the combined map shows 18 rather than 21 connections is because the three OE connections already exist in the E data. Setting the strength to ‘Both’ in the combined map shows 38 connections between these categories, whereas there are 35 in the E map and 10 in the OE map, and this is because there are three additional weak connections between categories in ‘1K’ and ‘1B’ in OE that are not found in E. There is going to be a lot of figuring out how exactly to amalgamate the data across all views and all levels, which is going to be a pretty large undertaking. But I feel that I’ve begun to figure it out, at least.
Week Beginning 20th April 2026
I spent much of Monday and Tuesday this week working on the DOST Auld Laws project, which I’d begun work on last week. I set up a test instance of the Cantaloupe IIIF image server on my laptop and began experimenting with the OpenSeadragon IIIF image viewer. It was pretty straightforward to set this up and to create a test page where the viewer connected to an image stored in the IIIF server and enables the user to zoom and pan around the image. The trickier issue was setting up the infrastructure so as to allow lines of text in the image to be clicked on and highlighted, and for IDs associated with each of these line regions to be accessible by JavaScript code (to eventually enable the corresponding line in the text view to also be highlighted).
Luca is currently working on a digital edition that allows regions of an image to be highlighted when buttons are pressed on and he very helpfully gave me access to the test site he’s working on so I could see how things work. I had been expecting that the region data would need to be somehow stored in the IIIF manifests for each image and that the IIIF server would be a lot of the processing to enable the display of and interaction with regions in the image, but both Luca’s site and another one I was referencing stored the region data as coordinates pulled into the front-end from a separate source, such as a JSON file. I was very happy to go with this approach, as firstly it meant I didn’t need to spend ages investigating IIIF manifests and how to update them and secondly it means the region data is not tied into a specific server technology, meaning it will be easier to repurpose it in future if that technology changes.
The line data exported from Transkribus consisted of detailed coordinates such as:
1309,2018 1469,2068 1716,2037 1876,2123 2092,2037 2234,2111 2389,2062 2555,2123 2734,2062 2796,2111 2975,2068 3555,2130 3629,2080 3808,2111 4062,2037 4457,2043 4524,2080 4599,2018 5055,2099 5142,2049 5321,2099 5438,2043 5586,2086 5845,2043 5901,2080 6271,2062 6271,1901 6129,1944 5802,1839 5710,1920 5284,1889 5234,1938 5166,1883 4771,1907 4481,1796 4333,1907 4222,1870 4043,1957 3950,1907 3808,1920 3753,1864 3574,1913 3506,1852 3426,1907 3006,1839 2895,1895 2438,1870 2265,1913 2055,1815 1975,1852 1759,1839 1660,1920 1506,1920 1309,184
But for the most part these refer to simple rectangles and with our data this level of detail isn’t really needed. Thankfully there is a script available to simplify the coordinates (https://github.com/dariok/page2tei/blob/master/simplify-coordinates.xsl) and I was able to adapt this to export the pixel coordinates of each corner of the line rectangle, such as (for the above polygon):
1309,1796 6271, 1796 6271, 2130 1309, 2130
OpenSeadragon doesn’t work directly with pixel dimensions, but has a function to convert these into its required format (see https://openseadragon.github.io/docs/OpenSeadragon.Viewport.html#imageToViewportRectangle) and I could therefore plug in the rectangle data and make the viewer display a region overlay on the image. I made a test that displayed all lines at once, as you can see below:
There is quite a lot of overlap between the lines here, but this isn’t a big issue. These areas will not be displayed all at once – only one will be displayed when the user clicks on the image. If the highlighted line isn’t the one the user wants due to there being an overlap it’s very easy for the user to just click again in a slightly different location to highlight the required line.
Having managed to get the regions displayed on the image, the next step was to process click events. Despite the overlays appearing as HTML elements, each sharing the same class, it was not possible to simply use jQuery to process clicks on this class. Instead, OpenSeadragon’s clickHandler needed to be used. With this in place it was then possible to process the click, to add a new highlighting class and to grab the ID of the element, which will eventually be used to find and highlight the line in the text view.
With all of this in place I then wrote a script to export line data (IDs, coordinates) from the TEI XML and store all of this in JSON files (one per document) that can then be loaded in whenever a page is displayed. I then moved on to setting up an initial version of the website for the digital edition, setting up an initial interface, menus, pages and such things. It’s all still running on my laptop for now, but I’ve made really good progress this week and will continue with it next week.
For the remainder of the week I worked on a variety of other projects. I made some further changes to the user survey popup for the Dictionaries of the Scots Language and I added some text to Sara Pons-Sanz’s AHRC proposal document. I set up the mapping metaphor site at the new URL we will be using for the site, meaning everything is ready for the major updates that I’ll hopefully be working on in the coming weeks.
I also wrote a document outlining the ‘facts and figures’ popup for the Burns Supper map. This took some time to research, but gives a handy overview of what the popup will contain in terms of summary data and visualisations.
Finally, I returned to the Place-names of Armagh project. The existing data for this project specified location data as Eastings and Northings and I need this as latitude and longitude. When I’d previously imported the data into QGIS the markers were all in the wrong place and I didn’t know why. Thankfully Frances Kane, who is working on the project suggested that this might be because I hadn’t set the coordinate reference system to Irish Grid, and this proved to be the answer.
After making the update the data all loaded at the correct locations and with this in place I was then able to follow this answer: https://gis.stackexchange.com/a/64700 to generate latitude and longitude fields for the data and then export this from QGIS as a CSV. After that I imported the data into the CMS and all records now have latitude and longitude. There’s still more I need to do, though. At the moment the map in the ‘edit place’ page in the CMS is generated based on the supplied grid reference or ITM coordinates, which the records still don’t have, so no map displays. I think I have a script that generates this data from latitude and longitude and I’ll look into this next week. I’ll also run the data through my script that grabs the altitude for places from Google Maps by sending latitude and longitude to the service. The front-end map uses latitude and longitude directly (no messing about with grid references or ITM) so I should be all set to start deploying the front-end map soon.
Week Beginning 23rd February 2026
I continued to work on the Burns Supper Map from Monday to Wednesday this week, with my first task being to import the public domain data into the map, taking the total number of suppers on the map to 751. I then implemented the advanced filter options. You can now open the advanced filter popup and select any of the filters you’re interested in. When you press the ‘apply filters’ button at the bottom of the popup it closes and the map displays only those suppers that match your criteria. The ‘Advanced’ section of the side menu then displays the number of matching suppers, your selected filters and buttons to refine or choose new filters, as the screenshot below demonstrates:
Note that I haven’t done any work on the colour schemes for the map yet – it’s all just using the colours taken from the Iona place-names map, but this will change. I also updated the map so that both the advanced filter and regular filter options now reposition the map to show all matching suppers when selected. However, icons on the left of the map can get obscured by the map menu. It has a tendency to sit on top of the US. I’m not sure what to do about this – I could ensure the map zooms out further, or I could close the side menu, although this might just confuse people.
I also updated the non-advanced filters section so that it displays the total number of suppers. However, a supper may have multiple filter options in a filter type so this total will not be the same figure as adding up all of the counts for individual options (e.g. one supper can have multiple toasts so will appear in the count for each individual toast that it features). I then created the table view. Pressing on the ‘Table view’ button will display all suppers currently found on the map in tabular form, and you can reorder the rows by pressing on a heading (e.g. ordering the rows by country). Pressing on the venue name link closes the table view, centres the map on the supper and opens the supper’s in-map record. Note that if you’ve performed a filter then the table only contains the filtered data.
Setting up the advanced filter option was especially time-consuming to implement so it was good to get that finished. There’s still a lot left to do, such as ensuring filter options get added to the page URL to enable bookmarking / sharing / citing of specific results. This is going to be another big job, and one that I’ll hopefully tackle next week.
I also spent some time this week drafting some text with Luca for a page about the technical developers across the College and the services we offer. This is not yet live, but will be a useful information point for staff who are looking to create an online resource for their data.
On Tuesday this week I participated in a meeting to discuss a new proposal being led by Sara Pons-Sanz at Cardiff that Glasgow will be involved with. It was a useful meeting and we all have a clearer idea about what the project will entail and how Glasgow will contribute. I had a further meeting on Friday with Jennifer Smith and Janine Illian about statistical analysis of the Speak For Yersel data and it looks like we’ll be getting some people in statistics working with our data, which is great.
I spent the remainder of the week working on the Playbills project. I’ve begun to set up an API and pages that will allow people to browse the playbills data. We still have a lot of work to do with the data, such as creating single, canonical records for venues, plays and other data types, so it’s likely that all of the data will need to be replaced at some point, but as the data structures are mostly finalised I decided to start work on some aspects of the front-end. So far I’ve created pages for browsing venues, listing playbills and browsing genres. I hope to continue with this next week.









