Week Beginning 31st August 2026

I took Monday and Friday off this week and on Wednesday and Thursday I attended the Digital Humanities Congress in Sheffield, meaning I only had one day to spend on regular work this week.  I spent most of this writing a specification document for the database and CMS for the Scottish Life Writing project, which I sent off to PI Craig and RA Emily for feedback.  I also had a further chat about the new search and results filtering options for the DSL.

I also spent a bit of time fixing a few bugs that I’d spotted in my recent updates to the HiMuJe Malabar map.  Category labels in the legend were getting duplicated when switching between menus in the map (e.g. ‘Area’ would become ‘AreaArea’), and I managed to fix this.  I’d also spotted that the categorisation type was getting removed from the URL when switching between menus, which meant that bookmarking or citing the page didn’t give the correct categorisation.  I managed to fix this too.  Finally, I’d spotted that loading a travel route stage from a URL into a fresh browser or browser tab meant that the map didn’t zoom to the stage’s location but instead displayed the full extent of the route.  This was because the code to position the map to the full extent was still running when the code to zoom to a specific location was triggered, meaning the latter did not take effect.  I fixed this by adding in a delay.  There is now a brief pause between the route loading in and then the map zooming to the chosen stage, which I think works quite well.

It was great to attend the conference in Sheffield and I attended a lot of interesting sessions (see https://www.dhi.ac.uk/dhc/2026/), a lot of which involved discussions of AI, which is having as huge impact on digital humanities research at the moment.  The opening plenary talk discussed ‘Literary Criticism after Large Language Models’.  The speaker noted that LLMs can produce a close reading of texts in seconds, but discussed whether this can actually ever be considered literary criticism or is it just imitation, and is there actually a difference?  Can it be considered criticism if there’s no understanding behind it?

The speaker gave an interesting overview of the stages of the use of computers in the humanities, starting with concordances and indexes in the 50s to the 90s, the emergence of distant reading in the 2000s, the shift to algorithmic criticism and interpretative visualisation in the 2010s and in the 2020s a major shift to using LLMs to interpret text – no longer just counting words and visualising but interpreting texts too.

The speaker pointed out the criticism involves interpretation (extracting meaning), surprise (new insights), context (historical and cultural), revision (deciding what is relevant) and reproducibility.  He discussed a study that compared how GPT-4 and humans can interpret metaphors, and how the AI was pretty effective, but whether this was from generalising or interpreting from existing knowledge.  Another study looked at authorship and genre in short passages and LLMs performed very well, with Llama 38b proving to be the best in this study.  The speaker then discussed a test he carried out on a poetry magazine corpus that he created and asked GPT-5 to analyse.  This was a new corpus so the LLM had no prior knowledge and consisted of around 1300 poems.  So this was ‘blind corpus analysis’ with no existing or contextual knowledge.  He asked the LLM to identify formal, stylistic, thematic, rhetorical and linguistic patterns to formulate a critical account of the collection.  The results were good, with genuine insights and the results could plausibly appear in literary scholarship.  He pointed out that LLMs can perform corpus level criticism but tend towards overcoherence, but can revise when challenged.  He also pointed out that everything the LLM does is heavily prompt dependent – the AI only ever does what is specifically asked – a human critic is still absolutely needed.

The speaker concluded by stating that criticism is a text about a text – it is an accountable act of reading and LLMs can do this.  He stated that there is ‘genuine’ (human) versus ‘false’ (AI) interpretation, but that the dividing line between the two is very thing now and wondered what forms of criticism will become more viable when plausible readings are cheap.

For the first parallel session I chose session 1, which involved three talks about AI.  The first looked at creating handwritten text recognition (HTR) for complex manuscripts using generative AI in the context of the Newton Project (https://newtonproject.ox.ac.uk/).  This has been around since 1998 and currently incudes around 12 million words of transcribed text.  They have a new project launching in 2027 which will digitise three books such as the Philosophae Naturalis Principia Mathematica.  These works have a lot of mathematics in them and are not uniform documents, including lots of hands, tables, diagrams, revisions, calculations and both Latin and English text.  The manuscripts are already digitised and are available in IIIF hosted at Cambridge Digital Library.  The speaker sampled 5% of the manuscript (around 80 pages) to evaluate the document using a 5-tier system.  19% were evaluated as being extremely difficult to transcribe with only 6% being easy.

The speaker had experimented with Transkribus and wanted to compare this to more general AI tools (Leo AI and Gemini).  She picked out four pages at different tiers to test, performing a line-by-line comparison of the tools.  For a tier-2 page (pretty easy) Leo performed best, followed by Gemini then Transkribus.  For a tier-4 page (difficult) Transkribus failed, Leo quit mid-way through and Gemini started to give summaries.  The speaker updated the prompt, which improved things but the results were still not highly accurate.  The speaker noted that as Transkribus works line by line it does not work so well whereas the AI tools take the whole page into consideration.

The speaker focussed on Gemini and created a detailed prompt with lots of instructions and developed a tool using Claude Sonnet that can be used to evaluated and correct HTR output (see https://johedesan.github.io/newt/).

The second speaker discussed marking up historical texts using AI.  He used data from the Proceedings of the Old Bailey project (https://www.oldbaileyonline.org/).  This was manually annotated using XML tags to note who did what, sentencing and other things.  The speaker wanted to investigate how AI could do this now, given prompts and examples.  He used the a locally installed instance of Qwen AI to perform the tests.  It did well, but there were some issues.  The output specified was not a conversational response, so there weren’t any hallucinations and he set it up with a ‘head’ using the spaCy NLP tool.  This was able to pick out defendants, victims and verdicts from trial documents.  He stated that the tools found assigning the correct verdict to be hard as this often appears at the end of the text and is tricky when there are multiple people.  Long trials are also tricky.

The third speaker discussed person identification in historical correspondence using an LLM.  The data was the Hammer-Purgstall correspondence (https://gams.uni-graz.at/hpe) which features 3545 transcribed letters.  This started with a print edition that was converted to TEI, although the semantic annotation is incomplete.  The project used the CorrespSearch webservice (https://correspsearch.net).

The project had to deal with many issues such as deciding on the boundaries for annotation (what should be included in the annotation) , dealing with tables, working out who pronouns refer to, and issues with families and groups.  Also people with mononyms, issues with different people with the same name (e.g. fathers and sons).  The project used an LLM as it could make use of context and works with multiple languages.  As the project didn’t want the LLM to mess with the XML the output was not updated XML but a series of positions where the relevant tokens start and end.

The results of the tests were compared to human annotation.  AI tests were carried out using GPT-5, Claude 5 and Qwen 3.5 and the tools were able to extract 68% of the names.  It took humans 13 hours to process the data but the AI tools only took 20-35 minutes.

For the next parallel session I chose Session 5, which was mostly focussed on GIS and mapping.  The first speaker discussed visualising the historical evolution of urban form using GIS.  The speaker used ArcGIS to analyse maps of colonial Calcutta over 100 years.  She looked at street sinuosity and street width to compare colonial centres to indigenous areas , and also ‘non-plot’ areas (areas in transition without formal buildings).  She created colour coded clusters based on these criteria (e.g. wide, straight, very low non-plot) to show the character of districts, comparing this to historical data and also showing how the situation changed through various different maps over 100 years.

The second speaker looked at land reclamation in 19th century Denmark based on historical documents that were available as PDFs.  The speaker used an LLM as it could work well with Danish.  This was able to extract coordinates and create georeferenced data that could then be searched.

The third speaker discussed geoparsing data found in archaeological records such as field reports – ‘grey literature’ that is otherwise difficult to extract data from.  The project developed a tool to work out which places are mentioned in the texts.  This is a two-part process involving recognition (finding places in the texts) and resolution (working out what the place actually is).  The latter needs the most work.  NER tools can struggle, e.g. when multiple places share the same name such as Memphis in the US and Memphis in Egypt.  The speaker discussed how the project’s tool creates a ‘bounding box’ to work out which places are likely correct based on the proximity of other named places in the text.  Further information can be found here: https://cloud.gate.ac.uk/.

Day 2 of the conference began with more parallel sessions, and I chose to attend Session 7.  The first speaker talked about the digitisation and study of marginalia.  He pointed out the marginalia were once frowned upon and attempts were sometimes made to remove them, such as through chemical erasure, but they are now of great interest.  He presented the work of a project called Magic (https://www.magic.unina.it/en) that is attempting to read these marginalia using modern technology.  The project uses FITS images that have embedded metadata and runs processes to reduce things like bleed-through to enhance the visibility of the notes – processes that also help with HTR.

The second speaker discussed the Law Versus Practice project (https://www.lawvpractice.ie/), which extracts data about women and property ownership in early modern Ireland.  It is creating biographical profiles of women from many primary sources, such as the list of claims, which contains 3,100 claims.  The data is fairly standardised, tabular data that is mostly printed, although there are inconsistencies, omissions and errors.  The speaker pointed out that disambiguation was an issue and needed research using titles, parentage, dates and geography.  This was tricky for women due to marriage and name changes.  The project has extracted over 5000 unique people that are linked to places, although the latter is often incomplete.

The third speaker discussed slavery in the Dutch East India Archive (see https://voices.iisg.nl/).  The project alanysed 7.5 million pages of transcriptions from the archive, mainly focussing on early modern Asia and mapped slave voyages.  HTR was carried out by a tool called Loghi (https://github.com/knaw-huc/loghi-htr) with linguistic analysis using AntConc.  The speaker noted that the transcriptions were not perfect but were good enough for searching and finding the co-occurrence of slave-related terms using AntConc.  Around 800 terms were analysed, including spelling variants.

The next session I attended as Session 10, and the first speaker was Jeremy Smith, who gave a fascinating talk about religious vocabulary in the 17th century, using the ‘quads’ concept from the Linguistic DNA project and data from EEBO-TCP.  He talked specifically about evidence from the (mostly religious) works of Lucy Hutchinson.  He discussed how the phrase ‘infinite love’ was not scriptural but became commonplace during this period.  He also discussed ‘saving faith’, a term which also appeared during this period.

The second speaker discussed how ideas travel in published works, mapping the influence of John Locke in the 18th century through direct quotation, paraphrase, concepts, critique, adaptation and appropriation.  The vocabulary can change so the speaker investigated connections using semantic similarity across 32 million pages in ECCO using 100-token chunks to investigate twenty high-use Locke quotes.  The third speaker discussed identifying essay-length republication in books and newspapers in the 18th century, looking at the work of David Hume.

For the final parallel session I attended Session 13.  The first speaker discussed news consumption, social media and youth engagement in Portugal.  The speaker pointed out how social media is now the gateway for news, either intentionally or incidentally, and how young people are more disengaged with other media.  There is no stable relationship with a specific new brand and engagement is formed through algorithmically curated feeds that shape an image of the world, including exposure to misinformation and propaganda.  The speaker looked at people aged 15-29 in Portugal using focus groups and questionnaires.  The data consisted of around 1500 responses, with more female than male participants, with most from the older age groups having higher education.  53% used social media for news daily, with TV news at 25% and newspapers 17%.  The respondents viewed traditional media as being more trustworthy, but social media also includes traditional news accounts as well as influencers and amateur news accounts.

The second speaker discussed propaganda and linguistic divergence in Russian online discourse during the war in Ukraine.  She discussed linguistic variation across state and social media, with social media more focussed on mobilisation and disinformation while state media focussed on pacifying the population and normalisation.

The third speaker gave an overview of preparations for a study of zeitgeist in UK parliamentary debates, looking at the period from Callaghan, Thatcher and Major (1976-1997).  The project will use the Hansard corpus.

The final session of the conference was a roundtable discussion about Digital Humanities in a changing higher education landscape and discussions covered the position of DH at a time when funding is increasingly hard to come by, and the role of AI in DH research.

The conference was hugely interesting and featured some great speakers, lots of interesting research and some interesting challenges.  It was really great to be able to attend.

Week Beginning 25th August 2026

This week I continued with the interactive map for the HiMuJe Malabar project and made some good progress with its development.  All subtypes now have an associated icon that is used on the map when the ‘type/subtype’ categorisation is selected, and I also reworked the colours of these, as you can see in the following screenshot:

After last week’s frustrations with the map legend, I reworked it to divide it a little better into sections by type, which you can also see in the above screenshot.  When a location popup is opened the corresponding subtype icon also now appears in the popup header.  I also added an option to the ‘Home’ menu to make all marker labels permanently visible on the map (until you turn them off again).  With all locations on this does clutter up the map, but I think it’s a useful option for people who are looking for a place but aren’t entirely sure where it is.  With labels on you can see the names of all places without having to hover over each marker in turn.  This works better when the map is filtered in the legend.  I also intentionally set the map so that I you set the labels to on and open the travel route menu the labels are removed as the travel routes would be very difficult to see with the labels on.

I also updated the colours for the other categorisation types to hopefully make them work a little better.  When the ‘number of references’ categorisation is selected the marker colours are now a gradient.  It’s also now possible to cite a specific stage in a travel route.  This appears in the citation text, e.g. “HiMuJe Malabar Digital Map: Watercolor map, travel routes menu, showing the travel route ‘Brahmins’, with unrelated places hidden, viewing stage 2 of the route, placenames categorised by language. 2026. In HiMuJe Malabar. Glasgow: University of Glasgow. Retrieved 28 August 2026 and following URL opens the route at Stage 2.  This option will mean I’ll be able to update the location popup to add a new tab that lists all of the stages in routes that the place appears in, and give links to each of these.  I haven’t done this yet, though.

Also this week I spent quite some time working for the Dictionaries of the Scots Language.  This included attending an online meeting to discuss the structure of the new site, at which some fairly major revisions to the new design were proposed.  I also engaged in a couple of email conversations about the updates that still need to be undertaken before the hopefully launch the new site later this year.  In preparation for all of this work I also went back through all of my emails from the DSL team and my notes to create a bit ‘to do’ list containing everything that needs done.  It’s going to be a busy couple of months.

Also this week I met with Craig Lamont and his RA Emily Hay to discuss the Life Writing project.  They’re going to be compiling a bibliography of Scottish life writing and I’m going to develop the systems for this.  I did look into using existing bibliography software but nothing I looked at did exactly what they are after so I’m going to create a simple CMS for them to use myself, with a front-end consisting of search and browse options and a map interface sometime next year.  I spent a bit of time after the meeting going through the sample spreadsheet they had prepared and sending a list of questions about the data.

I also met with Deven Parker this week to discuss an AHRC proposal she’s involved with.  I can’t say too much about this, but it’s related to the work we’re doing on the playbills and I would be involved in some technical capacity.  I also had a chat with Garrick Allen about a proposal he’s trying to get funding for after an earlier AHRC submission was unsuccessful.  I read through the documentation and gave feedback, and I’ve also put Garrick in touch with my colleague Luca Guariento who is probably going to provide the technical assistance.

Other tasks I carried out this week included investigating an issue with the Anglo-Norman Dictionary (which turned out to be an issue with the data and not the system), making some final tweaks to the Burns Supper Map before the project RA finished working on the project, and generating some Books and Borrowing data for a publication Katie Halsey is working on.

 

Week Beginning 17th August 2026

I continued to work on the DOST Auld Laws project on Monday and Tuesday this week.  My first task was to implement the tag selection for the Advanced Search.  As discussed last week, this section allows users to either include or exclude tags from their searches, or limit their search to only look at the contents of one of more specified tags.  Luca has already created the XML search options that would allow this, and my job was therefore to process the user’s selections, format the query and connect to the XML.  I managed to complete this on Monday and it’s now possible to construct a query such as “find all occurrences of words beginning ‘ȝe’ in ‘Edzell Doouments’ excluding any text found in ‘Expanded Forms’, Aitken Notes and Other Notes”.  It’s also possible to press on the ‘Refine your search’ button to return to the search form and the search criteria will be remembered in the form.  I also added in a ‘Clear’ button.

I continued working on the site on Tuesday.  I added a ‘Cite’ option to the search result page, which allows users to share or cite a particular page of a search, including the ordering.  I also added a help popup about the tag selection to the search form and I integrated the ‘show all Aitken notes’ query with the Advanced Search.  This is a special case that returns the full contents of all notes by Aitken across all or a selection of documents without the need to supply any search text.  Information about how to do this appears in the help popup as follows:

“To retrieve all Aitken Notes enter a dash (-) into the search box, select documents (if required), set Aitken Notes to Limit to content in this tag and press the search button.  This will display the full contents of all Aitken notes in your chosen documents.”

The Aitken Note results are a bit different to the regular results as they return the full contents of each tag, not a KWIC.  Therefore all of the text is highlighted in yellow and the sorting by left and right of the term doesn’t do anything, as there is nothing left or right.

Whilst working on the Advanced Search I’d spotted that the advanced search KWIC (KeyWord In Context) was not crossing line boundaries, meaning that when a work is found at the beginning of a line no contextual text is returned before this.  This is different to the Quick Search, which ignores line boundaries and is not ideal, as it means the search results are inconsistent and the quick search actually gives better results than the advanced search.  Thankfully Luca was able to find a solution to this and to make the advanced search results consistent with the quick search.

I’ve now pretty much done all I can do for this project until I hear back from Joanna, who is going to supply ancillary content and edited XML files.

I spent a lot of the rest of the week working on the HiMuJe Malabar project, continuing to develop the map.  On Thursday I attended a meeting with Ophira, Renu and project Co-I Ines, who is in Glasgow from Germany for a while to work on the project.  We had a very productive meeting and we now have a much clearer idea about how the travel routes and other aspects of the map will function.

In terms of actual development of the map, this week I fixed an issue with the travel routes whereby when multiple travel routes are selected, deselecting one would remove all locations, even those that were associated with another active route.  This took quite some time to sort out, unfortunately, but it’s fixed now and the routes work a lot better.

I then set about making it possible to share or cite the selected travel routes, also noting whether unrelated places have been hidden or not.  This required some major reworking of the code to ensure that these options are added to the URL and when the page loads that the options are taken from the URL and processed.  The citation text also needed to include information about the selections too, and any location selected when the route is active can also be cited, e.g.

“HiMuJe Malabar Digital Map: Satellite map, travel routes menu, showing the travel route ‘The migration of Jews and Christians in the Qissa’, with unrelated places hidden, placenames categorised by type / subtype, placename record for Mecca. 2026. In HiMuJe Malabar. Glasgow: University of Glasgow. Retrieved 19 August 2026”

At the meeting we decided what icons to use for the location categorisation so on Friday I returned to the map legend to rework this.  I started off by slightly reworking the data to ensure that all locations had both a type and a subtype, as a few locations only had the former.  This meant creating a new ‘Area’ type and making ‘Region’ a subtype of this, whereas previously ‘Region’ was the type.  The reason for doing this is so that the legend could be grouped by type, giving each type a count of the total number of locations contained across all subtypes and to include a checkbox for each type that when pressed on would select and deselect all subtypes.

That was the plan, anyway, but unfortunately I had a very frustrating day and didn’t manage to get this working by the end of it.  The difficulty is that Leaflet generates and processes the legend internally based on the map layers and trying to hook into this and change the default behaviour is very difficult.  I managed to get the ‘Type’ checkboxes to select and deselect all corresponding ‘Subtype’ checkboxes but doing so was not actually triggering any changes on the map – the corresponding map layers were not being affected when the checkboxes changed even though manually pressing on the checkboxes did trigger the layer change.  I’m afraid I ran out of time with this and didn’t manage to find a solution.  For now the types are just headings in the legend without associated checkboxes, which is a bit of a shame but I just don’t have the time at the moment to look into this further.

Also this week I made a few more tweaks to the data for the Burns Supper Map, including adding in a new batch of images, discussed migrating a couple of sites to our third-party hosting supplier, read a document ahead of next week’s DSL meeting and read some information Deven Parker had sent me about a funding bid.

 

Week Beginning 10th August 2026

I was back in Glasgow and back to a five-day working week this week, as the summer holiday period drew to a close.  I spent a fair amount of time this week continuing to work on the DOST Auld Laws project.  Luca has been working on the XML queries I’d specified for the advanced search and I was able to test them out and give feedback on them.  These are mostly all working as I’d hoped, which is really great, and I was able to begin working on the front-end aspects of the advanced search, with the aim of connecting all of this into the queries Luca had created.

However, I also had to spend quite a bit of time working with the source files, as the project PI Joanna Kopaczyk-McPherson had realised that two of the eight documents should really be split up into smaller sections.  This was no straightforward task, as not only did it require the XML documents to be split into smaller sections, but many other aspects needed to be updated.  The document IDs needed to be changed, which mean IDs used for image filenames throughout the documents also needed to be changed, with the filenames of the actual images also then needing to be updated too.  For page navigation in the site data is stored in a database and this also needed to be updated, and the XML files stored in eXist for search purposes also needed to be updated.  There were twelve steps I needed to follow for each required split, which took some time, but thankfully the process went pretty smoothly and we ended up with 13 documents in the site instead of the original 8.

There were further issues to come, as Joanna has now noticed that the documents need further edits and tweaks, and not just minor changes to text but structural issues such as the insertion of omitted lines.  This is going to be very difficult to do as lines are linked to coordinates in the images, which was all handled via the Transkribus tool.  We exported the XML files from Transkribus month ago and many major changes have been made to the files since the export so it’s not going to be possible to re-import them.  Manually creating new lines in the XML files as they are now will not have the connections through to coordinates in the corresponding image files, so we’ll end up with inconsistent data.  It’s not a great situation to be in, and ideally all editing of the documents should have been completed in Transkribus before we exported the files, something I’d mentioned back when I undertook the export process months ago.  We haven’t reached a decision on how best to handle this yet and we’ll continue to discuss the options next week.

Despite all of this I did manage to work on the advanced search, creating the advanced search form which, as specified, features a textbox where you can enter some text, a list of documents from which you have the option of selecting the ones you’re interested in and a list of tags that you can either include, exclude or limit your search to, as you can see in the following screenshot:

As of yet I have not implemented the tag limits, but the limit by documents is operational.  This connects through to Luca’s new eXist-db queries to perform a search limited by documents and as with the quick search, you can also sort the results by words to the left and right of the term, although I’ll need to get Luca to look into how the KWIC is generated for the advanced search results as they don’t seem to cross line boundaries, unlike the quick search.  The advanced search results display information about the documents you’ve selected and if you press on a search result to load the corresponding page the link back to the search results takes you to the right place.  There’s still lots to do.  The limit by tags is the biggest thing and will probably take quite some time.  I also still need to add in an option to cite a specific search result page and add in an option to refine your search, which will remember the options you previously filled in when you return to the search form.  An option to clear the search also needs to be added.  I’ll continue with this next week.

Also this week I participated in two meetings about the place-names AHRC proposal I’m involved with, and this is coming together very well.  We now have an outline proposal completed and pretty much ready for submission.

I also spent a bit of time continuing with the travel routes for the HiMuJe Malabar project, adding in a few more travel routes that had been prepared and made a few tweaks to the XSLT that generates the entry HTML for the new DSL interface, fixing some issues with the layout of the new ‘combs’ sections.

Week Beginning 3rd August 2026

I worked a total of four days over the past two weeks, and was on holiday for the remainder.  During this time I had a meeting to further discuss a place-names related AHRC proposal I’m involved with.  I can’t really say much more about it at this stage, but the proposal is coming together.  I also had to spend some time working with IT Support and Luca to figure out why our local server kept going offline repeatedly.  It looks like this was caused by the server getting swamped by requests from one particular source (almost certainly bot or AI) and thankfully IT Support were able to block this, after which the server was stable again.  It’s something we’re going to have to keep looking out for in future.  I also spent a bit of time working with Luca to get automatic WordPress updates working on our local server, as the way sites had been set up meant that the setting was not working.  Luca managed to find a solution to this, which is really great.

I also spent a bit more time working on the Burns Supper Map, creating a record for it on this site (see https://digital-humanities.glasgow.ac.uk/project/?id=156), adding more suppers that had been submitted via the survey and making some requested edits to existing suppers.  I also set up access to Google Analytics for the other two members of the project team.

In addition, I investigated an issue with the Scots Syntax Atlas after someone suggested that the linguists’ atlas was looking somewhat blurry.  I managed to figure out why this might be the case, although I’m not entirely sure whether this is a new issue or if the markers always looked like that.  I contacted the project PI and suggested a couple of updates, but I haven’t heard back yet so will need to wait and see what she says.

I spent most of the remainder of my time working on updates to our new test interface for Dictionaries of the Scots Language and working through the list of outstanding items for the DOST Auld Laws project.  For the DSL I completed the updates to the bibliography page that I began working on a couple of weeks ago.  I implemented pagination of the entries associated with bibliographical items, with navigation bars appearing above and below the entries, with 20 appearing per page and ‘jump to page’ buttons also appearing, just like with the search results.  This works pretty well, but is somewhat cumbersome for someone like Sir Walter Scott, who is referenced in 2861 entries, split over 144 pages.  We don’t have this issue with the search results are these are capped at 500 (25 pages) so we might need to think of other ways of handling this.

I also ensured that headword searches that don’t yield any results automatically perform a fulltext search for the term supplied.  This works on the live site, but only when both dictionaries are selected.  With the new site there are separate quick searches for SND and DOST so the additional search wasn’t being triggered.  It is now, as is the advanced headword search for both dictionaries.

I also made tweaks to the DSL’s new regional map based on feedback I’d received – adding in some content where we previously had placeholder text and ensuring the ‘About’ popup didn’t disappear off the bottom of smaller screens and a few other small updates.  I then began to look at the ancillary pages and how we can make them look a bit nicer.  I spent a bit of time on the ‘Word of the Week’ page and liaised with William Ashford, who is responsible for such content about this, and further updates that were going to make to the ancillary content closer to the launch date of the new site (which will hopefully be in November).

For the DOST Auld Laws project I added the top navigation bar that will link the site in with the SCOTS Corpus and CMSW.  I added in copyright information and added facilities to download page images and the XML files for each document.  These being up a pop-up asking for users to abide by the license before leading to the actual content, which hopefully won’t be too annoying.  I also added in the ‘cite’ popup to all document pages, which took a little time to implement, and added in Google Analytics.  I also made the image thumbnails on the document overview pages smaller and placed them in a collapsible section that is closed by default, plus I removed the introduction to the documents page, as this will be covered by the homepage.

I also removed some pages that didn’t have content (e.g. blank pages) from the beginning and end of some of the documents and I added a feature to turn off and on the line highlighting feature.  The highlighting feature allows the user to click on a line of text in the image or text for a page and for that line to be highlighted in both the image and the text, which is pretty nice.  Unfortunately the line highlighting gets in the way of the image viewer’s zoom and pan functionality on touchscreens, making it a somewhat unreliable and frustrating experience.  This new feature removes the option to ‘click’ on a line, meaning pointer events are not intercepted and make their way reliably through to the image viewer, which works much more smoothly.

Also this week I had an email conversation about user feedback and walkthough videos for the STAR resources and booked my accommodation for the DHC conference in Sheffield.  Next week I’m back in Glasgow and back working a full week, with summer holidays all over.

Week Beginning 20th July 2026

I worked on Thursday and Friday this week, having taken the other days off as a holiday.  Whilst I was away there was an issue with the server on which we host a lot of our important websites, meaning they were unavailable from Sunday afternoon until about 5pm on Tuesday, and I had to spend some time liaising with Luca about what should be done and responding to users and project members who were unable to access the sites.  I also had to spend some time once I was back updating everyone on the situation, and on Friday afternoon Luca and I met with Mike Irwin from IT Services to discuss the situation and what we could learn from it.  Hopefully we have a plan but we’ll just need to see how things work out in future.

Over the weekend the Burns Supper Map (https://burns-supper-map.gla.ac.uk) had its official launch (thankfully this website is not hosted on the server that encountered issues) at the Burns Birthplace Museum.  As I was away on holiday I wasn’t able to attend, but I was kept updated by the project team and it all went very well.  After the launch a few more suppers came in, and some existing suppers needed their details updating, for example because their title was not quite right, or their position on the map needed tweaking, or further photos were submitted.  I spent some time on Thursday making these updates.  I also exported the data for the Old English Thesaurus as CSV files, which are going to be submitted to the Oxford Text Archive, responded to some queries from the Dictionaries of the Scots Language team and made a minor tweak to the parts of speech section of entries on our test server.

Other than my meeting with Luca and Mike during the afternoon, I spent most of Friday working on the travel routes for the HiMuJe Malabar project.  I had been sent data for three travel routes to add to the map to test the travel route system I’d previously developed.

I updated my map code to link directly to the data source for the map that is generated from the digital edition and hosted at the University of Jena, so that new updates to places will automatically get pulled in.  I also updated the code so that when you select a route from the menu the map displays the full extent of the route, which I think will be very helpful.

Some of the places that appear in the travel itineraries are not yet found in the places file, or are found but do not yet have location data so for now I’ve had to remove these from the travel routes.  This isn’t a huge issue though, as for now the routes are really only for test purposes and I can add the missing locations in once they are available.

By the end of the week I’d added the three new routes to the map, although there is still a lot of work that needs to be done.  For example, the colours used for the markers and routes are still not finalised and we’ll definitely need to ensure the route colour is not also used for a polygon.  Currently Malabar is red, and so is the route, which makes things very confusing.  I also need to work on the menu to break it up into sections by type and assign different colours to the routes (or possibly the types).  But here is a screenshot showing the full extent of one of the trade routes:

I also had an email conversation with Tom Bartlett about a podcast site he created a while back that he would like to migrate to the University system.  I’m going to have a meeting with him about this, hopefully in the next few weeks.  I’m on holiday again next week and some of the following week, so it will be a while until my next update.

 

Week Beginning 13th July 2026

This was a four-day week for me as I’d taken Friday off (and I will also be off for the first three days of next week).  I finally managed to assign some time this week to implementing updates to the new Dictionaries of the Scots Language interface based on feedback that had been sent to me earlier this year.  I spent most of Monday and Tuesday working on this.

I can’t share any screenshots of the new interface at this stage, but I updated the search results box in the entry page to remove the tab for the other dictionary when performing a search (quick or advanced) for a specific dictionary.  This avoids misleading people as it otherwise the tab displays zero results for the other dictionary when in fact it just means the dictionary wasn’t actually searched.  Now when a quick search is performed, the other tab heading is replaced by a link to search the other dictionary.  Pressing on this performs whatever search you’ve executed (quick or advanced) on the other dictionary.

I also tweaked the font colour of the inactive tabs.  I realised that the white text on grey made it look like the tabs were disabled, when they’re not, they’re just inactive.  I therefore made the font darker, which I think works a lot better.  I then added in a button that scrolls the page to the search / browse box.  This appears above the entry header (and in the ‘sticky’ header that appears as you scroll the page) and only appears on narrower screens (where the infobox appears below rather than beside the entry text).  Where the entry is in the search results the text is ‘Scroll to results list’ with a down arrow.  Otherwise the text is ‘Scroll to browse list’.  The DSL team had requested that results term highlighting should be off by default, so I made this change too.

I then began to rework the bibliography page based on feedback.  We’ve decided to go with the version of the bibliography page that displays the quotations from any associated entries in addition to the headwords and links through to the entry pages.  This required some reworking of the API so that rather than returning individual citations, it brings back entries with each associated citation as part of this.  This means that multiple citations for an entry no longer appear as separate items in the list but are grouped by their entry, much like the quotation search results.  I also updated the count above the citations to display both the number of entries the item appears in as well as the number of citations (e.g. ‘Cited 718 times in 597 entries’) and each citation also includes its date as well now.  However, the order of citations needs to be the order they appear in the entry and not date order, otherwise the links through from the citation may end up taking you to the wrong one in the entry (as I discovered when I set things to date order).

The display of entries and citations is not exactly identical to the search results.  There is no sparkline as unfortunately this is not included as part of the bibliography data and I’d need to rework the database and API in order to include them.  Also, the quotations are in a larger font than the ones in the quotations search results as they seemed a bit small, and the headword isn’t highlighted in the quotations as this is something performed by the Solr search engine and isn’t available as things currently stand for the bibliographies.  I still need to add in pagination, which I didn’t have time to work on this week but will hopefully implement soon.

I also participated in a Teams meeting with the DSL this week that involved their interns reporting back about user engagement and observations about the website.  It was very interesting to hear their feedback and it will give us lots to think about as we continue to improve the resource.

Other than working for the DSL, I also made some last-minute updates to the Burns Supper Map before its official launch over the weekend.  The resource is now available for anyone to use at https://burns-supper-map.gla.ac.uk.  I wasn’t able to attend the launch as I was away on holiday but from what I’ve heard it was a great success.

I also spent a bit of time fixing an issue with the Anglo-Norman Dictionary where certain entries that had an apostrophe in their headwords (e.g. j’) were not loading while others were.  The reason for the discrepancy was because some headwords had curly apostrophes and other had straight ones.  The straight ones were getting encoded (e.g. “j'”) and were then not getting found.  Once I managed to figure this out I was able to fix the issue.

On Wednesday I had a lengthy online meeting for the HiMuJe Malabar project to discuss the interactive travel routes.  The meeting lasted about two and a half hours but it was worthwhile as we all have a much clearer idea of how to proceed with the routes now.  I just need to wait until the team sends me some initial travel routes using the spreadsheet template I sent them and then I’ll be able to continue my work on this.

On Thursday I participated in a Teams call with colleagues from Nottingham and Cardiff Universities to discuss a place-names proposal that I am likely to be involved with.  I can’t really say much more about it at the moment, but it’s all sounding very interesting.  I spent a few hours after the meeting writing a document containing some initial thoughts about how the technical infrastructure for the project could work.

I was then off on holiday on Friday and I won’t be back at work again until next Thursday.

Week Beginning 6th July 2026

I returned to work on Monday this week after a lovely holiday.  Most of my time this week was spent working for the Dictionaries of the Scots Language, processing a new dataset exported from their editing system and integrating it into the website.  The reason this process took considerably longer than usual is that the data had a major structural difference:  Entry parts of speech had been moved, rationalised and restructured.  Previously parts of speech appeared as elements within <meta> but also appeared embedded in the main entry where they were only tagged with HTML italic tags.  There were more than 500 different combinations of parts of speech and things like homonym numbers were mixed in with them too.

The DSL editors have spent a huge amount of time working on the parts of speech to separate them out, standardise them and ensure other data such as homonym numbers are stored in a separate but related manner.  The new structure also ensures that entry parts of speech are only stored once in the entry XML using a structure that makes sense semantically rather than for display only.

As the new entry XML in the data export now differed markedly from the earlier structure I then had to rewrite my data processing scripts, update the database and Solr cores, and the API and front-end to deal with the new structure.  This was a lot of work.  My first step was to create a new table to hold the individual POS data, consisting of a part pf speech and associated ‘hom’, ‘syn’ and ‘infl’ values.  I then updated my extraction script so that the existing pos field in the entry table now gets its content from the ‘origpos’ attribute (so we continue to have a record of the original part of speech) and then to populate the new POS table to store each individual pos, including the value, hom, syn and infl data (where applicable) for each entry.

With this update in place I could then (after running a few smaller-scale tests) process every entry in the SND and DOST export files, which resulted in 35,655 parts of speech records being generated for SND entries and 49,828 being generated for DOST.

My next step was to update the entry browse order, which is used to decide in which order the entries appear in the browse pane when viewing entries.  This previously used the old POS system to decide in which order entries with the same headword appeared (e.g. so that nouns appeared first).  I had to update this to use the new POS system, and also ensure that the POS labels, which are used in the browse pane, the search results and the entry header.

With this update in place I then ran my scripts to process the citations and bibliographies.  After that I worked on the Solr cores we use for search purposes.  These consist of an ‘entry’ core used for headword and fulltext searching, and a ‘quotation’ core used for searching the individual quotations.  Both of these needed to be updated to include new POS fields, which will eventually allow me to add parts of speech filters to the search facilities.  I created new Solr cores for each, both featuring new POS fields, after which I would update my script that generates the data to populate the cores to ensure the POS data was included.

With the new cores successfully populated with the new data I then needed to update the API to work with the new POS fields (e.g. so that the labels returned for use in the search results and browse pane use the new POS data).  After that I then needed to update the XSLT scripts that process the entry XML files for display to ensure that the new POS structure is displayed when viewing an entry.

At this point all of the updates were still running on my laptop, and with everything in place and tested the final step was to migrate everything to our test server.  I don’t have direct access to the server so needed the help of someone with this access in order to complete tasks such as creating and populating the new Solr cores, and thankfully Luca agreed to help out.  The process went remarkably smoothly and by the end of Thursday the update was complete, with our test server displaying the new POS fields.  I also made a number of other minor updates to the display of entries via the XSLT files that had been requested and it’s now over to the DSL editors to test everything out and make sure all is working as it should.

I spent the remainder of the week dealing with emails I’d received whilst I was away.  I also had a meeting with Joanna about the DOST Auld Laws project, which is nearing completion now, and had a meeting with Renu regarding the display of polygons in the map for the HiMuJe Malabar project.

Week Beginning 8th June 2026

I began feeling unwell on Monday this week, but managed to struggle through until Thursday morning, by which point I just couldn’t sit at my desk any more.  I was then off work sick on Thursday and Friday.

Despite coming down with something I still managed to get quite a bit done on Monday to Wednesday this week.  On Monday I mainly focussed on the Playbills project.  I manually sorted all but three of the remaining plays that didn’t have canonical records and asked Deven to further investigate the remaining three.  I then added the canonical play data to the output produced by the API and wrote a script that would output all of the playbill data for every playbill as JSON files.  This will be used for offline analysis by Deven.

On Tuesday I continued to work on the map for the HiMuJe Malabar project, mostly tidying up some loose ends with the map.  I implemented the URL shortener and set up the ‘share / cite’ options so that now if you open a place record and press on the ‘Share / Cite’ tab this loads content, with the citation text being dynamically generated based on the map options that are currently selected.  It’s also now possible to press the ‘Share’ button at the bottom of the map menu to share or cite a specific map view as opposed to a specific record.

I also implemented the ‘Reset map’ option on the ‘Home’ menu and pressing on this resets the map to the default view and categorisation type and I finally sorted out the polygons on the map so they should always now appear behind the markers, even when turning layers on and off in the legend.

The last thing I implemented was the ‘Table view’.  If you press on this button at the bottom of the map menu it opens a popup containing all of the data in tabular form.  I haven’t included the references in this as this is too much data for such a table, but I have included a count of the number of references for each place.  Similarly, I’ve only included the latitude and longitude for each place and not the full GeoJSON shapes for the polygons.  You can press on a column heading to order the table by that column (press a second time to reverse the order).  I think it will be quite a useful view, for example if you know the name of a place but are not sure exactly where it is located you can order the table by placename and find it.  Pressing on a placename in the table closes the popup, centres the map on the location and (after an intentional brief delay so you can appreciate where on the map you’re looking at) the relevant record popup opens.  I might see about highlighting the place’s marker to make it clearer which is the relevant one in areas where there are many.

I still need to implement categorisation by certainty, but as of yet we don’t have any certainty data in the JSON file so I’m leaving this for now.  I’m intending to start on the travel routes soon.  We don’t have any routes available yet, but I’m going to create a test route and create the interface for plotting the connections along the route so we can see how this might work.

On Wednesday I continued to work on the DOST Auld Laws project.  I managed to fix the issue with multiple results in one line causing the in-page navigation to stop working.  This required quite a bit of reworking as the results were identified by line and instead I needed to add in a new way of identifying that a result may be on the same line but is actually a different result.  I did this by counting the number of results per line and using this figure in addition to the line ID to track the results, adding this counter to the URL that is used to reach the document page from the search results, with the document page then taking this figure and using it to ascertain which result is the current one and allowing the results traversal on the document page to function.

I then moved onto looking at results page ordering.  Previously the search results were ordered by document name, document page and line, but I wanted to add further sort options via a drop-down list to the right of the ‘you searched for…’ box.  In addition to the default ordering, this allows you to order by highlighted term (useful for searches like ‘*other’ where the start of terms may differ) and a concordance-like word position.  The latter allows you to sort the results by up to 5 words to the left or right of the term, with the selected position highlighted in cyan.  The screenshot below shows the search for ‘*other’ ordered by the second word to the left of the term:

 

This was a particularly large and complex update to implement as many parts of the code had to be rewritten, but I think it’s worth it.  The final update I implemented was to ensure that the ‘return to search results’ button in the search bar of the document page now takes you back to the specific results page and retains the ordering you’ve selected rather than taking you to the first page of results and the default ordering.

Also this week I set up a bare-bones WordPress site for Craig Lamont’s new project, and I’ll meet with him over the summer to get this fully set up.  I also tweaked the Burns Supper Map to ensure long subtitles for the charts don’t end up overlapping with the ‘hamburger’ menus for each chart and read through the findings of the user survey for the Dictionaries of the Scots Language that I’d been sent, and which was on the whole very positive.

Week Beginning 1st June 2026

The project I spent the most time working on this week was the DOST Auld Laws project, which I hadn’t worked on for a couple of weeks.  The last time I worked on the project I’d managed to get an initial version of the quick search working, but there were some issues with it.  A phrase search wasn’t working, wildcards at the start of a search term were not working, the full term was not getting highlighted in the results where an <expan> is present in the term, I hadn’t implemented the pagination of search results and I still needed to update the page view to enable results traversal directly from the page when accessed via the search results.

I was struggling somewhat with phrase searching and full term highlighting and ended up going round in circles with ChatGPT for several hours and getting nowhere.  I eventually asked my colleague Luca, who has considerably more experience and knowledge of eXist-db and XML document querying, for some advice.  Thankfully he was able to come up with a query script that did exactly what I needed it to do, which was a massive help as I was really struggling.  This is definitely an example of where a conversing with actual person is much more effective than AI and  I really must give Luca credit for the help he gave.

By the end of the week my updates meant that it was possible to use wildcards at the start of a search term, as you can see from the following screenshot that shows the results for ‘*other’:

With Luca’s help, phrase searches now work, and the following screenshot shows the results of a search for “the landis of”:

Again with Luca’s help, the search term in each result is now fully highlighted even when an <expan> is present in the word, so for example a search for ‘pebillis’ previously found two results but only highlighted ‘pebill’ in the results as the term was recorded as ‘pebill<expan>is</expan>’.  But now the full term is highlighted.

I also implemented results pagination, with is currently set to display a maximum of 20 results per page.  The following screenshot shows a search for ‘r?cht’ with the pagination in place:

And now when you reach a document page that is in the search results (e.g. by selecting it from the search results) a search results navigation bar appears above the document navigation bar, allowing you to navigate to the next or previous search result or return to the full search results, as the following screenshot demonstrates:

I’ll probably add in a ‘clear search results’ button here too, and I still need to fix an issue with the results navigation when there are multiple results on one line.  At the moment there is no way for the code to differentiate these so the ‘next’ and ‘previous’ links get stuck.  I also need to look into speed issues with the search too.  But some good progress has been made this week and I feel much more confident using eXist-db now.

Also this week I spent a bit more time on the Burns Supper Map, adding in videos and images for the recently imported supper records.  I also had an email conversation with the Dictionaries of the Scots Language people about parts of speech and new data imports, which will probably be taking place in the next few weeks.  I also attended a ‘coffee and catch-up’ with the other developers in the College, which was really valuable as always.

The rest of my week was divided between the Playbills project and the HiMuJe Malabar project.  For Playbills I processed a spreadsheet containing around 400 performers that Deven had manually processed last week.  I ended up doing a bit more manual tweaking after investigating the appearance of some of the performers in the playbill images, and when I ran the spreadsheet through my import script we then had 52,178 performers in the system, and of these some 52,048 of these have a gender assigned, which I pretty amazing.

I then moved onto processing the plays based on the spreadsheet Deven had worked on before Easter that notes which performances actually feature the same play, even if it is not referred to in exactly the same way.  Using this I created canonical records for plays, I set some marked plays as ‘special attractions’  and I deleted certain plays that had been marked for deletion.  Before I did this I updated the database so that all tables include an ‘isactive’ field, and updated the API so that it only includes data where ‘isactive’ is set to ‘Y’.  Then when it came to deleting plays I didn’t actually delete them but set them (and all associated data such as performers and roles) to ‘isactive = N’.  This means if we realise something has been ‘deleted’ that needs to be reinstated I’ll just need to update the relevant fields back to ‘isactive = Y’ rather than having to find and re-insert all associated data.

For the most part the script was successful and result in 261 plays being deleted and 1098 converted to special attractions. It then created 2039 canonical records and 4134 other plays were then assigned to these.  There were a few issues with some plays referencing canonical records for other plays that were set to be deleted (so canonical records weren’t created for them), and I passed these on to Deven so she could look into them.

I then created an API endpoint and front-end pages for listing the canonical plays and the details for a selected canonical play.  Below is a screenshot of part of the list of canonical plays, ordered by number of plays:

It’s just an initial version that lists the titles, the record type (either play or ‘special attraction’) and a count of the number of associated plays.  You can press on column headings to order the table by the column and if you press on a canonical play name you can access a list of plays that are associated with it, as you can see below:

This lists each associated play’s name, date, playbill, venue, location and genres, and you can click on each linked item to reach the relevant page (e.g. the details for a play or the associated playbill page).  This is just a work in progress and we’ll probably want to update it, for example to include a genre filter on the canonical plays list, or including access to lists of associated roles and performers.  I also updated the playbill page to add links through from play titles to the relevant play page, and I’ve added a ‘Type’ field for each play that either displays ‘Play’ or ‘Special Attraction’.

For the Malabar project I sorted out the legend for source texts, alphabetising the list and ensuring the pane has a maximum width.  I also added in the two other categorisation options that I’d included in my specification document:  Number of references and placename languages.  Number of references categorises the markers and polygons based on the number of times each place is referenced in the source texts.  I set up the categories to match the available data (0, 1-5, 6-10, 11-20, 21+) but these can very easily be altered as the data and number of references grow.  I think it will prove quite useful to be able to quickly identify the places that are referenced the most in the texts.  Below is a screenshot showing the currently available data categorised by number of references:

The other new categorisation option was ‘Placename languages’  and this categorises the places based on the languages of the placename variants included in each record.  This will likely need some further work, such as adding in full language names rather than the codes, and possibly filtering out ‘en’ as I’m guessing these would not have been found in the original sources, but it’s still interesting to use the categorisation – for example finding all placenames that have a Hebrew form.  Here is a screenshot showing placename languages:

The data itself is still being compiled and I still need to work on the marker colours (and icons) and also to ensure the polygons always appear behind the markers and don’t make the markers unclickable, as is sometimes the case at the moment.  Also this week I added the selection of base map, menu section and categorisation type to the URL, meaning it’s now possible to bookmark / share / cite specific views of the map.  You can also link directly to a specific record as well.