Week Beginning 1st December 2025
My time this week was divided between several different projects and meetings. I spent quite a bit of time working with Transkribus, ahead of Friday’s Transkribus workshop at which I was speaking about my text extraction experiments with the Edinburgh Gazetteer (https://edinburghgazetteer.glasgow.ac.uk/). Back when I worked on the project with Rhona Brown (almost ten years ago now) we’d attempted to extract the text from the images using OCR but our experiments had been pretty hopeless. I attended a Transkribus event in Edinburgh earlier this year and had done a little bit of work with the Gazetteer in Transkribus then, but hadn’t progressed very far. This week I made considerably more progress, experimenting with a three-column subsection of one page, as you can see below:
Transkribus can identify columns of text by analysing what it calls ‘Fields’ so this is what I asked it to do first, using the ‘Baroness of Blocks’ model. This is something that’s only available with a subscription, but I was able to make use of a free trial. Unfortunately the process was not very successful. It did correctly identify the columns, but split the image up into sections within each column, with some parts of the image entirely missing from the classification (e.g. the top of column 2) in the image below:
I manually corrected this using the interface, as the following image shows:
However, I’m uncertain how I would be able to train the system to automatically and correctly identify such columns for other images – this would require further investigation. The next step was to identify lines within each column, which is accomplished using a ‘Layout’ model. I chose the default ‘Mixed line orientation’ model which was pretty successful in identifying all of the lines in each column. It wasn’t perfect but it was good enough for test purposes, as you can see below:
The third step was to extract the text. For test purposes I wanted to see how the model would work without any training, and I chose the ‘Text Titan I ter’ model. This took several minutes to process, but the results were very encouraging, as the following image demonstrates:
There were some issues, however, such as the large drop-characters at the beginning of sections being omitted, and some words that are legible to humans being incorrect, such as ‘acie’ instead of ‘acre’. The ends of lines in the first column were also missing, so line identification would need to be tweaked. Despite these issues the text is broadly understandable and complete. The poor print quality and the long ‘S’ character were processed successfully and some sections that were very difficult for a human to decipher were processed successfully by the tool. There are still issues to be ironed out with regards to successfully identifying columns of text, and these would need to be addressed before any batch processing of the entire Gazetteer, but it’s looking very promising.
I spent most of Tuesday this week attending the presentations for the new Grade 8 and Grade 7 roles for the post-graduate course in Digital Humanities that is being set up in Information Studies. It was really interesting to hear the presentations and to learn more about what the candidates would bring to the roles. There were some really excellent candidates and it was very useful to hear from them.
Also this week I spent a little time working on the Place-names of Ayrshire project ahead of next week’s launch and engaged in a continuing email conversation about how the data for the interactive map will be gathered and stored for Ophira Gamliel’s Malabar project. We’ve now managed to reach an agreement on how to proceed with this, which is a relief. I also spent some time continuing to make updates to the Place-names of Armagh resource, creating a nice interface for the project website using suggested public domain images and fonts, and adding parishes and baronies to the content management system. The new project website is not yet live, but here’s how the new design currently looks:
I’ll meet with the project team next week to give a run-through of the CMS and working with WordPress, after which they should be in a position to start adding data to the resource. I also managed to spent a little more time working with the data for the Fife place-names project, continuing to rationalise the place-name elements and their connections, but there is still more to do for this.
On Thursday this week I met with Deven Parker to discuss the requirements for an online resource for her Playbills project. We discussed the kinds of search and browse facilities she would like to include and other features such as visualisations and data summaries. Next week I’m going to write up the requirements and share the document with her. Also on Thursday I met with Wendy Anderson and Carole Hough to discuss some potential future updates to the Mapping Metaphor resource. I can’t really go into any details here, but we had a good meeting and will consider the options before meeting again in the New Year to see how this might be taken forward.





