Week Beginning 14th July 2025

It’s been a while since my last update due to holidays and conferences.  The week beginning the 23rd of June I worked on Monday and Tuesday and then I was on holiday on Wednesday the 25th of June and didn’t return to work until Friday the 11th of July.  Most of these three days were spent finalising my presentation for the DH2025 conference, going through the sizable conference programme to decide which of the parallel sessions to attend and sorting out arrangements for travelling to, from and around Lisbon.

I also made comments and suggestions for a document the DSL had prepared about the process for exporting data from their editing system into the online database and updated the copyright information for the recently updated IPA chart videos on the Seeing Speech site.  As all videos previously had the same copyright statement, this was hard coded into the video overlays, so I needed to update the database in include flags for different statements and the code to then process the statements depending on the flags.

On Sunday the 13th I flew to Lisbon and attended the DH2025 conference (https://dh2025.adho.org/browse-the-program-agenda/).  I didn’t attend any of the pre-conference workshops, but instead spent come of Monday exploring the city and finding the venue, and the rest of the time working: doing some last-minute run-throughs of my paper, making further updates to Seeing Speech, as it turned out some of the ExtIPA videos also needed their copyright statements updated.

On Tuesday afternoon the main conference began with a plenary, which discussed AI and the use of Large Language Models and Vision Language Models (LLMs and VLMs) in DH research, and also discussed storing data in a graph database (neo4j) and using the cypher graph query language, and getting AI to use this.

My paper was scheduled for the first session on Wednesday, which was great as it meant I could get it over with and enjoy the rest of the conference without having to worry about it.  The presentation went well and I covered everything in the available time without any issues.  However, the server that hosts Speak For Yersel and other resources went offline before my talk, which was really bad timing and rather embarrassing.  Thankfully I wasn’t doing a live demonstration, but I had to tell everyone during my talk that they couldn’t view the resource as it was offline.  I frantically tried to get people to sort the issue out before my talk began, but no-one was around at that time.  Our IT people got the server back online again by the end of the session, but it was too late by then.

There were a few questions afterwards, none of them particularly awkward.  They all related to class – I think people were particularly interested in how the education level of our participants was markedly different from the census data.  One person wanted to know whether you could compare the results of people with different education levels, and I was able to discuss the filter options which offer this facility.  Another person wondered whether there was anything that could be done to increase engagement with groups from lower education levels.  I didn’t really have an answer for this, but talked about how targeted face-to-face approaches like SCOSYA are perhaps more effective at engaging directly with such groups.  Another person suggested that we could in some way use the geographical data we have about our participants to tie this into the census data for these areas to gain further insights into their likely class and backgrounds based on where they live.  I thought this sounded like a really interesting idea, even if it would obviously lead to assumptions being made about participants.  Something to think about.

The other papers in my session were also generally about map-based resources – one used ArcGIS to map treasure-hunting expeditions, another mapped connections between cultural-heritage research, scholars and outputs.  Another project discussed chopping up historical maps into squares and extracting features found in these squares, for example finding patches that feature railway lines to analyse the populations that existed in close proximity to railways, and more information can be found here: https://data.nls.uk/data/map-spatial-data/living-with-machines-railspace-building/.

The second session I attended also focussed on mapping.  One paper looked at extracting places from historical Japanese prints, using AI tools such as SigLip (https://huggingface.co/docs/transformers/en/model_doc/siglip) to identify features such as boats in images and then using crowdsourcing to identify real-world locations found in the prints via a ‘street view’ style interface.  Another paper extracted places mentioned in the ‘Baltische Briefe’, a German-language newspaper from the Baltic states.  This used Named-entity recognition (NER) to extract places from the text using the SpaCy tool (https://spacy.io/).  Another paper discussed the representations of Colombian communities in New York and London.  The speaker used the Mapbox Storytelling tool (https://labs.mapbox.com/storytelling/) for her presentation, which looks like a really great way of telling a story via maps.  The final paper discussed issues relating to Soviet nuclear testing in central Asia, which was both interesting and horrifying to learn about.

The third session I attended featured three papers.  The first was about extracting data from millions of French census records (see https://socface.site.ined.fr/en/ and https://socface.teklia.com/).  The project is extracting data from the entire French census from 1836 to 1936 – between 20 and 30 million pages.  The data is mostly tabular  and clean, but obviously handwritten.  It uses HTR (handwritten text recognition) and NER using YOLOv8 (https://docs.ultralytics.com/models/yolov8/) and DAN (https://gitlab.teklia.com/atr/dan).  The speaker compared the current project which uses AI approaches to an earlier crowdsourcing project and demonstrated how cheaper and quicker the AI approach is – it would appear that crowdsourcing for these types of projects has had its day.

The second paper also focussed on AI tools, using GPT-4o to refine the prompts given to the AI tools to improve the retrieval while the third used AI tools to extract and analyse scenes from images in German children’s books from 1800 to 1940, looking to identify scenes of play, reading, teaching and such things whilst identifying the genders of the participants.  Images were analysed using Doc-UFCN (https://pypi.org/project/doc-ufcn/) to detect and extract the illustrations from the pages and images were then classified using SmolVML (https://huggingface.co/blog/smolvlm), Llava (https://huggingface.co/docs/transformers/en/model_doc/llava) and Qwen (https://huggingface.co/Qwen).  They also used the Collection Space Navigator (https://collection-space-navigator.github.io/) to identify clusters.

For the last session of the day I didn’t find one that entirely appealed to me, and I attended the session on ‘Networks, Lexicons, Text Mining and Digital Philology’.  The first paper was presented in Italian (but with slides in English), so it was a bit hard to follow, and none of the other papers were especially relevant to anything I do.  One speaker did present some interesting visualisations using violin plots (see https://r-graph-gallery.com/violin.html) which I hadn’t seen before, so that was good to see.

On day two, the first session I attended discussed topic modelling on a corpus of poetry using NLTK.  The second paper looked at gender portrayal in Chinese preschool children’s books from 2012 to 2022.  The speaker identified 5922 books aimed at the 0-9 age group and used ChatGPT to identify categories, such as genders, race, age, and working out whether the characters were central or side characters.  The third paper discussed making TEI resources multilingual and the fourth explored gender differences in gaming culture.  This looked at game streaming on Twitch, and specifically the comments posted on male vs. female streamers’ videos.  It focussed on German language videos, with data collected for one week, looking at the top 16 female and male streamers.  370,000 messages were captured, but only 165,000 had more than 4 tokens.  Topic modelling was performed on these using BERTopic (https://maartengr.github.io/BERTopic/index.html) and clustered using HDBScan (https://hdbscan.readthedocs.io/en/latest/).  It was an interesting paper but there’s more refinement that could be done – the gender of message posters was not included, the genres of the games were not considered, so different audiences would likely be targeted.  The final paper discussed the ‘dark sides of DH’ and discussed how DH datasets can be full of bias, such as colonial and gender.

The second session I attended was about handwritten text recognition and AI.  The first speaker discussed ‘The Delineator’, a US women’s magazine from the late 19th and early 20th centuries.  The speaker discussed using Newspaper Navigator (https://github.com/LibraryOfCongress/newspaper-navigator) to extract images and used CLIP (https://huggingface.co/docs/transformers/en/model_doc/clip) to classify them.  The speaker also mentioned Chroma DB (https://github.com/chroma-core/chroma) as a means of storing data that can then be queried by AI tools.

The second speaker discussed using ChatGPT to generate SKOS thesaurus structures from images of the required data structures, either hand-drawn of created digitally, using both a real thesaurus and a fictional one.  The tool was able to generate the required structures with a high level of accuracy.

The third speaker discussed a tool to automatically transcribe Catalan manuscripts from the middle-ages.  These were notaries (wills etc), and the project included 750,000 images.  These included cursive handwriting, medieval Latin and large numbers of abbreviations.  They looked at a sample of 100 charters (3369 lines, 80 hands, 29 document types) from 1208 to 1499.  They manually annotated the 100 images using eScriptorium (https://escriptorium.rich.ru.nl/) to segment the lines which were then transcribed using the Kraken tool (https://hal.science/hal-04936936v1/file/0673a.pdf).

The fourth speaker looked at HTR models for 16th and 17th century Spanish writing (see the project website https://wp.lancs.ac.uk/newspainfleets).  The project is looking at ship registers for voyages from Spain to Mexico – around 10,000 documents with 4 script types.  They are developing HTR models for these using Transkribus.  In the training data, lines were manually identified to cover large ascenders and descenders and the project intend to publish their models as open access.  The speaker noted that image quality makes a big difference, with less and 1 megapixel being very bad, 5 being good and at anything over 8 the picture quality stops being an issue.

The final speaker discussed using LLMs to perform post-OCR error correction on historical French texts.  The speaker used Llama 3.2-3b instruct (https://huggingface.co/meta-llama/Llama-3.2-3B-Instruct) using prompts telling the LLM that it is performing OCR correction and to correct errors whilst preserving the 19th century style and retaining the line breaks.  The speaker noted that the output was worse than the original OCR, with semantic errors, formatting errors and context misrepresentation.  It was still interesting to hear about, though, and someone suggested giving the LLM the image as well as the OCR might help in future.

When it came to the third session of the day I was again uncertain which to attend, as none of them seemed especially relevant to me.  I ended up attending a session on ‘Text Mining, Tracing and Quantification in Literature’.  The first paper was about gender depictions in medieval Chinese epitaph verses and wasn’t really my kind of thing.  The second was about Lady Gregory’s Irish Legends and how she adapted the original works to support Irish nationalism, using NLTK and word2vec.  The final speaker examined references to Greco-Roman authors in modern academic discourse, looking at 56,116 articles across 16 disciplines from 1990 to 2019, analysed using Spearman rank correlation and Zipf distribution and generating UMAP (uniform manifold approximation and projection) plots to show clusters.

The final session of the day was on evaluating projects and tools.  This included a speaker who evaluated DH websites to find common objectives and to investigate whether these had been met and a speaker also reviewed DH websites, this time their interfaces and whether accessibility guidelines had been adhered to.  A further speaker discussed the creation of an AI assistant for the Basque language using retrieval augmented generation (RAG) to ensure the use of up to date information sources.  Another speaker evaluated the use of Graph RAG (https://microsoft.github.io/graphrag/) to retrieve semantically structured data, while the final speaker discussed creating apps in ChatGPT that could then be used to perform small tasks, such as converting TXT files to CSV.

I began the third and final day of the conference by attending a session on ‘mapping and visualising conflict, violence and slavery’.  The first presentation was about a large-scale historical and archaeological study of Basel in Switzerland, which is producing more than 20 published volumes and a lot of online data.  The paper was focussed on working collaboratively across different disciplines and the conflicts that can arrive between participants.

The second speaker discussed interactive maps of police violence in the US, which used hexagonal grids to represent the data rather than relying on boundaries that are created by authority and don’t reflect the real world.  Maps can be divided into triangles, squares or hexagons and the latter have 6 neighbours, each of equal distance, so work best.  Clusters were generated using Local Moran’s I (https://en.wikipedia.org/wiki/Moran%27s_I).  The speaker also discussed adding in a temporal element too, for example using ArcGIS space time cubes (https://www.esri.com/arcgis-blog/products/arcgis-pro/announcements/introducing-a-new-space-time-cube-visualization-experience-in-arcgis-pro).  The speaker sourced his data from https://mappingpoliceviolence.org/.

The third speaker discussed visualising resistance in the archive of slavery, and discussed a history of data visualisation from 1786 to 1900, and how the 19th century was the beginning of modern visualisation, discussing the work of William Playfair (https://en.wikipedia.org/wiki/William_Playfair).  She pointed out that all visualisations have an agenda, going on to show a contemporary visualisation of the layout of a slave ship and then demonstrating so beautiful but quite difficult to interpret visualisations about slave ship voyages, which can be accessed here: https://dataxdesign.io/chapters/description#voyage-interactive

 

The second session of the day also looked at visualisations, along with automating text processing using LLMs.  The first speaker looked at LLMs producing code, and a pipeline that could be used for code generation and refinement using a prompt from a human that is then passed through the pipeline, passing the prompt to GPT-4o, then refining and executing.

The second speaker discussed Pandore (https://obtic.sorbonne-universite.fr/developpement/toolbox/), a toolbox that can be used to perform OCR, format conversion, NER, topic modelling without the user needing any IT knowledge.  The third speaker presented ‘Flow filter’, a generalisable visualisation and query toolkit.  The speaker gave some excellent demonstrations of the tool, for example from the Saltaire census data, but I’m unable to find any information about the tool online to link to.

The fourth speaker presented about open science literacy and the final speaker discussed using LLMs for NER in Urdu, highlighting how many standard libraries do not work well with languages that are written right to left.

In the third session of the day there wasn’t anything that seemed hugely relevant, so I ended up attending a session that included a discussion of open archaeology in Catalan, a speaker who created ‘time maps’ for plays / novels / films that plot the actual passage of time on one axis and the order of events on the other axis.  The speaker used Pulp Fiction as an example, visualising how the various story segments take place in the film and how they are chronologically ordered.  The final speaker discussed automated vs manual subset selection in the Finnish national bibliography, looking at first editions for adults that were published between 1909 and 1917.

I also didn’t find anything in the final session that seemed all that relevant and attended a session that featured three papers on quite different subjects.  The first looked at how to measure dramatic texts, looking at key components such as plot, character, dialogue and action.  The speaker discussed the work of Boris Yarkho, a Russian literary scholar who looked at speech distribution in five-act tragedies, and also discussed vectorising plays – reducing them to an ordered sequence of numbers based on features.  The speaker made use of the DraCor corpus (https://dracor.org).  The second speaker was about Ukrainian Epigraphy, and more specifically about mapping resources and terms that had been created for Greco-Roman epigraphy so that they would be applicable to Ukrainian epigraphy.  The final speaker discussed how the post-colonial canon is dominated by a few Western-based English Language writers such as Salman Rushdie, and discussed whether broadening this out with translations of texts from South Asia could improve the situation.

The conference ended with the customary closing ceremony that involved some prizes being given out, some closing speeches and some information about future DH conferences.  DH2026 will be held in Daejeon, South Korea while DH2027 will be held in Galway, Republic of Ireland.  DH2028 will be held in Cape Town, South Africa.