Category: Thesaurus of Old English
Week Beginning 20th July 2026
I worked on Thursday and Friday this week, having taken the other days off as a holiday. Whilst I was away there was an issue with the server on which we host a lot of our important websites, meaning they were unavailable from Sunday afternoon until about 5pm on Tuesday, and I had to spend some time liaising with Luca about what should be done and responding to users and project members who were unable to access the sites. I also had to spend some time once I was back updating everyone on the situation, and on Friday afternoon Luca and I met with Mike Irwin from IT Services to discuss the situation and what we could learn from it. Hopefully we have a plan but we’ll just need to see how things work out in future.
Over the weekend the Burns Supper Map (https://burns-supper-map.gla.ac.uk) had its official launch (thankfully this website is not hosted on the server that encountered issues) at the Burns Birthplace Museum. As I was away on holiday I wasn’t able to attend, but I was kept updated by the project team and it all went very well. After the launch a few more suppers came in, and some existing suppers needed their details updating, for example because their title was not quite right, or their position on the map needed tweaking, or further photos were submitted. I spent some time on Thursday making these updates. I also exported the data for the Old English Thesaurus as CSV files, which are going to be submitted to the Oxford Text Archive, responded to some queries from the Dictionaries of the Scots Language team and made a minor tweak to the parts of speech section of entries on our test server.
Other than my meeting with Luca and Mike during the afternoon, I spent most of Friday working on the travel routes for the HiMuJe Malabar project. I had been sent data for three travel routes to add to the map to test the travel route system I’d previously developed.
I updated my map code to link directly to the data source for the map that is generated from the digital edition and hosted at the University of Jena, so that new updates to places will automatically get pulled in. I also updated the code so that when you select a route from the menu the map displays the full extent of the route, which I think will be very helpful.
Some of the places that appear in the travel itineraries are not yet found in the places file, or are found but do not yet have location data so for now I’ve had to remove these from the travel routes. This isn’t a huge issue though, as for now the routes are really only for test purposes and I can add the missing locations in once they are available.
By the end of the week I’d added the three new routes to the map, although there is still a lot of work that needs to be done. For example, the colours used for the markers and routes are still not finalised and we’ll definitely need to ensure the route colour is not also used for a polygon. Currently Malabar is red, and so is the route, which makes things very confusing. I also need to work on the menu to break it up into sections by type and assign different colours to the routes (or possibly the types). But here is a screenshot showing the full extent of one of the trade routes:
I also had an email conversation with Tom Bartlett about a podcast site he created a while back that he would like to migrate to the University system. I’m going to have a meeting with him about this, hopefully in the next few weeks. I’m on holiday again next week and some of the following week, so it will be a while until my next update.
Week Beginning 21st July 2025
Monday was a public holiday this week, which was much appreciated as I’d spent much of the preceding day travelling back from the DH conference in Lisbon. I spent the majority of Tuesday writing my report from the conference, processing my expenses and catching up with emails, which I continued to do throughout the week.
My two main tasks of the week were to add new data to the VARICS lookup feature and the begin the development of an online exhibition for the Robert Fergusson project. The former involved extracting data about speech measurements from a somewhat complicated spreadsheet and updating the data relating to the specific measurements in the online database. I also needed to create new methods for calculating outcomes for some of the new measurements, as previously all calculations were based on the inputted measurement being below 1.5 times the standard deviation whereas some new measurements needed to be above instead, and a further one needed to be within a specified range. I also had to add in some new measurement labels and other information.
For the Robert Fergusson project I’d been given access to a large collection of files that had been used for a physical exhibition about the poet at the Mitchell Library last year, and it’s my job to translate these into an online exhibition. Project RA Amy Wilcockson had arranged the data into folders for each section, with a Word file giving an overview of how each section will fit together, which was all very helpful. It still took a while to familiarise myself with the data, read through the various documents and begin to understand how the information might be presented online. By the end of the week I’d created the overall structure of the exhibition, by which I mean placeholder pages and navigation options for accessing the various pages, and I’d added in the content for the introduction and the first exhibition section, which is about Fergusson’s publications. So far the content is either text or captioned images so it’s not that complicated – more a matter of laying everything out nicely.
One thing I wanted to experiment was a means of animating the content into place as the user scrolls the page, a feature that makes the page seem more alive and interesting. What I didn’t know was how to achieve this, as it’s not something that is easy to search for online. I eventually found a JavaScript library called GSAP (https://gsap.com/) that is a very powerful tool for animating any aspect of a web page. It features a plugin called ScrollTrigger (https://gsap.com/docs/v3/Plugins/ScrollTrigger/) that (as you can probably guess) triggers animations as the user scrolls the page. I spent some time working with the library, but while the tool is obviously very powerful and flexible, I didn’t find the documentation to be all that useful. What I was really hoping for was a series of simple tutorials and demos showing all of the various effects, but the ‘demo’ page for the tool just features a series of more advanced and not especially broadly applicable interfaces rather than things like ‘here’s an example of a block of text swooshing in from the left’. The plugin page itself didn’t feature any working examples and mostly linked to a 20 minute YouTube video, which I personally don’t find as helpful as a text-based tutorial with working examples linked to.
Although I got the tool working and managed to get a couple of experiments set up I felt that my use case (make some blocks of text animate into place as the page is scrolled) wasn’t really covered very well and that I should perhaps look elsewhere. I then found a rather old page (from 2017) that discussed exactly the sort of behaviour I was hoping for (https://css-tricks.com/aos-css-driven-scroll-animation-library/) and found the library that was discussed was still operational (see https://michalsnik.github.io/aos/). This library was really simple and did exactly what I was looking for, so I integrated it with my exhibition pages, making sections animate into place as the user scrolls down the page. Of course I’ll need to see whether the project team likes this feature, but whether we use it or not I feel like I’ve learned something new and hopefully reusable by looking into this.
Also this week I had discussions with the DSL people about how the online site selects headword forms from the XML, and also about some significant changes to their data structures which will have implications for how the files are processed for and displayed on the website. I will need to rewrite my script that processes the XML files to handle the new structure to ensure the new or newly positioned data is extracted as it should be. I may need to update the database structure as well, in order to store things like the regional information, and it’s possible the Solr index structure may also need to be updated. The XSLT file that transforms the XML files into HTML whenever an entry is loaded will also need to be rewritten to deal with the new structure. So there will be a lot to do in the coming months!
I also fixed an editing error that had occurred with the Thesaurus of Old English and had an email discussion with Tony Harris about the Bilingual Thesaurus code I’d previously sent him for use on a new project, and how I process data and develop user interfaces. It looks like we’ll be meeting up in Zurich in January, which will be useful.
Week Beginning 26th May 2025
Monday this week was a public holiday, and I divided my four working days across several projects. The Iona place-names project is having its official launch on the 9th of June and a fairly last minute request for the map interface was to incorporate ‘thematic maps’ – maps that show a set of markers that share a common theme, such as bird names. Last Friday Sofia sent me the data for two example thematic maps (which I think we’re going to be calling ‘Virtual Trails’ in the public interface) and this week I set about creating the new feature.
When creating a new feature I’d usually create the sections of the CMS for managing the data before updating the front-end but as time is rather tight I thought it best to create the necessary structures in the database, the required updates to the API and work directly on the front-end with the sample data. If there’s time before the launch I’ll create the CMS pages, but if not and we want to add more maps I’ll just have to get the team to send me spreadsheets and I’ll add the data directly to the database. Below is a screenshot showing the new map:
I’m still working on this and there is still quite a bit to do, as adding in the new section and new map type has meant a lot of reworking of existing structures. But as the above screenshot demonstrates, the interface now features a new ‘Virtual Trails’ menu section that when expanded has a placeholder description and a ‘Choose a virtual trail’ button. Pressing on this opens a popup that lists the thematic maps, showing the titles and descriptions. Pressing on a title closes the popup, adds the selected map’s title and description to the left-hand menu and loads the relevant map markers into the map. You can then change the display options (e.g. turning labels to always on, changing the classification or base map) and open the records as you would with any other map.
So that’s the basics in place, but there are still many updates I’ll need to make to the front-end to fully integrate the new feature, including ensuring that the ‘Reset map’ resets the virtual trail menu contents, ensuring citations and bookmarks / sharing URLs work when a trail is selected, ensuring the table view works properly and cancelling out any already active search or browse options when a trail is selected.
Also this week I applied the new updates I’d made last week to the element glossary to the other place-name maps, such as Berwickshire. I also added ‘deselect’ to the legend, as apparently users were unaware that the ‘select all’ option could be used to ‘deselect all’ when unchecked., and I spotted that the Iona full map ‘cite’ option was referencing the Ayr site and fixed that too.
For Speak For Yersel, I spent some time this week creating the slides for my presentation at the Lisbon DH conference, and I have now completed a first version of the slides and script. I’m probably going to have to trim it down a little, though, as when running through it I was taking longer than my allotted ten minutes. I’ll probably have to take out the parts about how the maps were generated, which is a shame as it’s the most DH section, but it’s also not something we discussed in the abstract so if anything needs cut it’s the obvious choice.
On Thursday this week I met with Ophira Gamliel to discuss the interviews for her project, which are taking place next week, and in the afternoon I attended a Modernities Research Cluster event which featured two fascinating speakers. I also fixed an issue with the Thesaurus of Old English for Jane Roberts and fixed a problem with the batch update of citations in the Anglo-Norman Dictionary. This feature of the content management system allows citations across all entries (or selected entries) to be updated simply by editing the citation once. However, there was a problem when adding dates to citations that did not already have them. The script was running into problems when adding a date for an attestation that didn’t already have a <dateInfo> element. I’d included a check for this, but the check was causing a fatal error rather than executing the code that I’d written to deal with such attestations. Thankfully once identified it was relatively easy to fix the problem, and now the batch update system is working perfectly.
Week Beginning 18th November 2024
I spent a fair amount of time this week working on a paper about Speak For Yersel and its new regions for the DH2025 conference. I worked on this with the PI Jennifer Smith and after several revisions and a meeting on Tuesday we had something that was ready to submit. I also spent some time finalising the code for the tool that I created to generate the new survey areas. We’ve decided to make this publicly available for others to download and reuse via GitHub but before I could add the code I needed to get it ready for publication – doing things like ensuring libraries like Bootstrap and jQuery are referenced externally rather than including them directly, ensuring no passwords and keys are included, updating the setup instructions document, preparing the example data and running through a few test installations to make sure the thing actually works. For sample data we’re making the Wales data available. This consists of not only the questions and answer options for the survey but all of the geolocational data that is needed in order to ensure submitted answers can be correctly pinned onto the map. You can now access all of the code for SFY here: https://github.com/baitken/sfy.
Also this week I spent some time going through the website mockups and sample data that I’d been sent for the VARICS project. They are creating a database of speech files with filters, similar in many respects to the facilities I created for the STAR project and I’ll be developing a resource for the project over the coming months. We have a project meeting scheduled in a couple of weeks, and I have a lot of questions to ask when we meet.
I heard some good news this week regarding project funding. The Anglo-Norman Dictionary has had its final phase approved by the AHRC, after a few weeks of uncertainty, which is fantastic. Also Joanna Kopaczyk’s DOST project has received funding and I’ll be working with her on this in the new year, which is great.
Other tasks this week included updating the embedded Scots School Dictionary in Speak For Yersel to replace this with the new Concise Scots Dictionary that is now available from the DSL, plus giving some advice to the DSL editor William Ashford about which versions of the ancillary pages of the DSL site and its test equivalent to update. I also updated the links from the Thesaurus of Old English to the Dictionary of Old English as their URLs have changed, and discussed the creation of a new place-names site for Nairn with Thomas Clancy, submitting the request to have a subdomain created for this.
Week Beginning 8th January 2024
This was my first week back after the Christmas holidays and after catching up with emails I spent the best part of two days fixing the content management system of one of the resources that had been migrated during the end of last year. The Saints Places resource (https://saintsplaces.gla.ac.uk/) is not one I created but I’ve taken on responsibility for it due to my involvement with other place-names resources. The front-end was migrated by Luca and was working perfectly, but he hadn’t touched the CMS, which is understandable given that the project launched more than ten years ago. However, I was contacted during the holidays by one of the project team who said that the resource is still regularly updated and therefore I needed to get the CMS up and running again. This required updates to database query calls and session management and it took quite some time to update and test everything. I also lost an hour or so with a script that was failing to initiate a session, even though the session start code looked identical to other scripts that worked. It turned out that this was due to the character encoding of the script, which had been set to UTF-8 BOM, which meant that hidden characters were being outputted to the browser by PHP before the session was instantiated, which then made the session fail. Thankfully once I realised this it was straightforward to convert the script from UTF-8 BOM to regular UTF-8, which solved the problem.
With this unexpected task out of the way I then returned to my work on the new map interface for the Place-names of Iona project, working through the ‘to do’ list I’d created after our last project meeting just before Christmas. I updated the map legend filter list to add in a ‘select all’ option. This took some time to implement but I think it will be really useful. You can now deselect the ‘select all’ to be left with an empty map, allowing you to start adding in the data you’re interested in rather than having to manually remove all of the uninteresting categories. You can also reselect ‘select all’ to add everything back in again.
I did a bit of work on the altitude search, making it possible to search for an altitude of zero (either on its own or with a range starting at zero such as ‘0-10’). This was not previously working as zero was being treated as empty, meaning the search didn’t run. I’ve also fixed an issue with the display of place-names with a zero altitude – previously these displayed an altitude of ‘nullm’ but they now display ‘0m’. I also updated the altitude filter groups to make them more fine-grained and updated the colours to make them more varied rather than the shades of green we previously had. Now 0-24m is a sandy yellow, 25-49, is light green, 50-74m is dark green, 75-99 is brown and anything over 99 is dark grey (currently no matching data).
I also made the satellite view the default map tileset, with the previous default moved to third in the list and labelled ‘Relief’. This proved to be trickier to update than I thought it would be (e.g. pressing the ‘reset map’ button was still loading the old default even though it shouldn’t have) but I managed to get it sorted. I also updated the map popups so they have a white background and a blue header to match the look of the full record and removed all references to Landranger maps in the popup as these were not relevant. Below is a screenshot showing these changes:
I then moved onto the development of the elements glossary, which I completed this week. This can now be accessed from the ‘Element glossary’ menu item and opens in a pop-up the same as the advanced search and the record. By default elements across all languages are loaded but you can select a specific language from the drop-down list. It’s also possible to cite or bookmark a specific view of the glossary, which will load the map with the glossary open at the required place.
I’ve tried to make better use of space than similar pages on the old place-names sites by using three columns. The place-name elements are links and pressing on one performs a search for the element in question. I also updated the full record popup to link the elements listed in it to the search results. I had intended to link to the glossary rather than the search results, which is what happens in the other place-names sites, but I thought it would be more useful and less confusing to link directly to the search results instead. Below is a screenshot showing the glossary open and displaying elements in Scottish Standard English:
I also think I’ve sorted out the issue with in-record links not working as they should in Chrome and other issues involving bar characters. I’ve done quite a bit of testing with Chrome and all seems fine to me, but I’ll need to wait an see if other members of the team encounter any issues. I also added in the ‘translation’ field to the popup and full record, although there are only a few records that currently have this field populated, relabelled the historical OS maps and fixed a bug in the CMS that was resulting in multiple ampersands being generated when an ampersand was used in certain fields.
My final update for the project this week was to change the historical forms in the full record to hide the source information by default. You now need to press a ‘show sources’ checkbox above the historical forms to turn these on. I think having the sources turned off really helps to make the historical forms easier to understand.
I also spent a bit of time this week on the Books and Borrowing project, including participating in a project team Zoom call on Monday. I had thought that we’d be ready for a final cache generation and the launch of the full website this week, but the team are still making final tweaks to the data and this had therefore been pushed back to Wednesday next week. But this week I updated the ‘genre through time’ visualisation as it turned out that the query that returned the number of borrowing records per genre per year wasn’t quite right and this was giving somewhat inflated figures, which I managed to resolve. I also created records for the first volume of the Leighton Library Minute Books. There will be three such volumes in total, all of which will feature digitised images only (no transcriptions). I processed the images and generated page records for the first volume and will tackle the other two once the images are ready.
Also this week I made a few visual tweaks to the Erskine project website (https://erskine.glasgow.ac.uk/) and I fixed a misplaced map marker in the Place-names of Berwickshire resource (https://berwickshire-placenames.glasgow.ac.uk/). For some reason the longitude was incorrect for the place-name, even though the latitude was fine, which resulted in the marker displaying in Wales. I also fixed a couple of issues with the Old English Thesaurus for Jane Roberts and responded to a query from Jennifer Smith regarding the Speak For Yersel resource.
Finally, I investigated an issue with the Anglo-Norman Dictionary. An entry was displaying what appeared to be an erroneous first date so I investigated what was going on. The earliest date for the entry was being generated from this attestation:
<attestation id="C-e055cdb1"><dateInfo> <text_date post="1390" pre="1314" cert="">1390-1412</text_date> <ms_date post="1400" pre="1449" cert="">s.xv<sup>1</sup></ms_date> </dateInfo> <quotation>luy donantz aussi congié et eleccion d’estudier en divinitee ou en loy canoun a son plesir, et ce le plus favorablement a cause de nous</quotation> <reference><source siglum="Lett_and_Pet" target=""><loc>412.19</loc></source></reference> </attestation>
Specifically the text date:
<text_date post="1390" pre="1314" cert="">1390-1412</text_date>
This particular attestation was being picked as the earliest due to a typo in the ‘pre’ date which is 1314 when it should be 1412. Where there is a range of dates the code generates a single year at the midpoint that is used as a hidden first date for ordering purposes (this was agreed upon back when we were first adding in first dates of attestation). The code to do this subtracts the ‘post’ date from the ‘pre’ date, divides this in two and then adds it to the ‘post’ date, which finds the middle point. With the typo the code therefore subtracts 1390 from 1314, giving -76. This is divided in two giving -38. This is then added onto the ‘post’ date of 1390, which gives 1352. 1352 is the earliest date for any of the entry’s attestations and therefore the earliest display date is set to ‘1390-1412’. Fixing the typo in the XML and processing the file would therefore rectify the issue.
Week Beginning 4th December 2023
After spending much of my time over the past three weeks adding genre to the Books and Borrowing project I turned my attention to other projects for most of this week. One of my main tasks was to go through the feedback from the Dictionaries of the Scots Language people regarding the new date and quotation searches I’d developed back in September. There was quite a lot to go through, fixing bugs and updating the functionality and layout of the new features. This included fixing a bug with the full text Boolean search, which was querying the headword field rather than the full text and changing the way quotation search ranking works. Previously quotation search results were ranked by the percentage of matching quotes, and if this was the same then the entry with the largest number of quotes would appear higher. Unfortunately this meant that entries with only one quote ended up ranked higher than entries with large numbers of quotes, not all of which contained the term. I updated this so that the algorithm now counts the number of matching quotes and ranks primarily on this, only using the percentage of matching quotes when two entries have the same number of matching quotes. So now a quotation search for ‘dreich’ ranks what are hopefully the most important entries first.
I also updated the display of dates in quotations to make them bold and updated the CSV download option to limit the number of fields that get returned. I also noticed that when a quotation search exceeded the maximum number of allowed results (e.g. ‘heid’) it was returning no results due to a bug in the code, which I fixed. I also fixed a bug that was stopping wildcards in quick searches from working as intended and fixed an issue with the question mark wildcard in the advanced headword search.
I then made updates to the layout of the advanced search page, including adding placeholder ‘YYYY’ text to the year boxes, adding a warning about the date range when dates provided are beyond the scope of the dictionaries and overhauling the search help layout down the right of the search form. The help text scroll down/up was always a bit clunky so I’ve replaced it with what I think is a neater version. You can see this, and the year warning in the following screenshot:
I also tweaked the layout of the search results page, including updating the way the information about what was search for is displayed, moving some text to a tooltip, moving the ‘hide snippets’ option to the top menu bar and ensuring the warning that is displayed when too many results are returned appears directly above the results. You can see all of this in the following screenshot:
I then moved onto updates to the sparklines. The team decided they wanted the gap length between attestations to be increased from 25 to 50 years. This would mean individual narrow lines would then be grouped into thicker blocks. They also wanted the SND sparkline to extend to 2005, whereas previously it was cut off at 2000 (with any attestations after this point given the year 2000 in the visualisation). These updates required me to make changes to the scripts that generate the Solr data and to then regenerate the data and import it into Solr. This took some time to develop and process, and currently the results are only running on my laptop as it’s likely the team will want further changes made to the data. The following screenshot shows a sparkline when the gap length was set to 25 years:
And the following screenshot shows the same sparkline with the gap length set to 50 years:
I also updated the dates that are displayed in an entry beside the sparkline to include the full dates of attestation as found in the sparkline tooltip rather than just displaying the first and last dates of attestation.
I completed going through the feedback and making updates on Wednesday and now I need to want and see whether further updates are required before we go live with the new date and quotation search facilities.
I spent the rest of the week working on various projects. I made a small tweak to remove an erroneous category from the Old English Thesaurus and dealt with a few data issues for the Books and Borrowing project too, including generating spreadsheets of data for checking (e.g. list of all of the distinct borrower titles) and then making updates to the online database after these spreadsheets had been checked. I also fixed a bug with the genre search, which was joining multiple genre selections with Boolean AND when it should have been joining them with Boolean OR.
I also returned to working for the Anglo-Norman Dictionary. This included updating the XSLT so that legiturs in variant lists displayed properly (see ‘la noitement (l. l’anoitement))’ here: https://anglo-norman.net/entry/anoitement). Whilst sorting this out I noticed that some entries would appear to have multiple ‘active’ records in the database – a situation that should not have happened. After spotting this I did some frantic investigation to understand what was going on. Thankfully it turned out that the issue has only affected 23 entries, with all but two of them having two active records. I’m not sure what happened with ‘bland’ to result in 36 active records, and ‘anoitement’ with 9, but I figured out a way to resolve the issue and ensure it doesn’t happen again in future. I updated the script that publishes holding area entries to ensure any existing ‘active’ records are removed when the new record is published. Previously the script was only dealing with one ‘active’ entry (as that is all there should have been), which I think may have been how the issue cropped up. In future the duplicate issue will rectify itself whenever one of the records with duplicate active records is edited – at the point of publication all existing ‘active’ records will be moved to the ‘history’ table.
Also for the AND this week I updated the DTD to ensure that superscript text is allowed in commentaries. I also removed the embedded Twitter feed from the homepage as it looks like this facility has been permanently removed by Twitter / X. I’ve also tweaked the logo on narrow screens so it doesn’t display so large, which should make the site better to use on mobile phones and I fixed an issue with the entry proofreader which was referencing an older version of jQuery that no longer existed. I also fixed the dictionary’s ‘browse up’ facility, which had broken.
I also found some time to return to working on the new map interface for the Iona place-names project and have now added in the full record details. When you press on a marker to open the popup there is now a ‘View full record’ button. Pressing on this opens an overlay the same as the ‘Advanced search’ that contains all of the information about the record, in the same way as the record page on the other place-name resources. This is divided into a tab for general information and another for historical forms as you can see from the following screenshot:
Finally this week I kept project teams updated on another server move that took place overnight on Thursday. This resulted in downtime for all affected websites, but all was working again the next morning. I needed to go through all of the websites to ensure they were working as intended after the move, and thankfully all was well.
Week Beginning 1st May 2023
Monday was a holiday this week, so ordinarily this would have been a four-day week for me. However, I was unfortunately picked for jury duty and I was obliged to attend court on Wednesday and Friday, making it a two-day week. I’m also going to have to attend court on Tuesday next week as well (after next Monday’s coronation holiday) but hopefully that will be an end to the disruption.
On Tuesday this week I spent a bit of time working on the migration of sites to external hosting and spent the remainder of the day adding the new MRI 2 recordings to the IPA chart on the Speech Star website. I uploaded all of the videos and added in a new ‘MRI 2’ video type. I then uploaded and integrated all of the metadata. It took quite a long time to get all of this working (pretty much all day), adding the data to all four of the IPA charts, but I got it all done. I will need to update the charts on the Seeing Speech website too once everyone is happy with how the charts look.
On Thursday I made some further tweaks to the Edinburgh’s Enlightenment map and migrated three further sites to external hosting. I also spent some time updating the shared spreadsheet we’re using to keep track of the Arts websites, adding in contact details for all of the sites I’m responsible for and making a note of the sites I’ve migrated.
I also made some tweaks to the Speech Star feedback pages I’d created last week, populated a few pages of the Speech Star website with content from Seeing Speech, added content to the ‘contact us’ page, fixed some broken links that people had spotted in the site, swapped a couple of video files around that needed fixed in the charts and added explanatory text to the extIPA chart page. I also added in some new symbols to the IPA charts for sounds that were not present on the original versions but we now have videos for in the MRI 2 data.
I also investigated a strange issue that Jane Roberts had encountered when adding works to the Old English Thesaurus using the CMS. Certain combinations of characters in the ‘notes’ field were getting blocked by Apache, and once I’d figured this out we were able to address the issue.
I also spent a bit of time on the Books and Borrowing project, running a query and generating data about all of the book holding records that currently have no associated book edition record in the system (there are about 10,000 such records). We had also received the images for the final two registers in the Advocates Library from the NLS digitisation unit and I spent some time downloading these, processing the images to remove blank pages and update the filenames, uploading the images to our server and then running a script to generate register and page records for each page in both registers. These should be the last registers that need to get added to the system so it’s something of a milestone.
Week Beginning 30th January 2023
This was a four-day week as the latest round of UCU strike action began on Wednesday. Strike action if going to continue for the next two months, which is going to have a major impact on what I can achieve each week.
I spent almost all of this week working on the Books and Borrowing project. This first two days were mainly spent dealing with data related issues. This included writing a script to merge duplicate editions based on a spreadsheet of editions that I’d previously sent Matt to which he had added a column to denote which duplicate should be merged with which. It took quite some time to write the script due to having to deal with associated book works and authors. Some of the duplicates that were to be deleted had book work associations whilst the edition to keep didn’t. These cases had to be checked for and the book work association had to be transferred over.
Authors were a little more complicated as both the duplicate to be deleted and the one to keep may have multiple associated authors. If the duplicate edition to keep had no authors but the one to be deleted did then each of these had to be associated with the edition to keep. But if both the edition to delete and the one to keep had authors only those authors from the ‘to delete’ edition that were not already represented in the ‘to keep’ edition’s author list had to be associated. In such cases where an author did need to be associated with the ’to keep’ edition I also added in a further check to ensure the author being associated didn’t have the same name (but different ID) as one already associated, as there are duplicate authors in the system.
With all of this done the script then had to reassign the holding records from the ‘to delete’ edition to the ‘to keep’ one and then finally delete the relevant edition. As the script makes significant changes to the data I first ran it on a version of the data I had running on my laptop to check that the script worked as intended, which thankfully it did. After completing the test I then (after taking another backup of the database in case of problems) ran the script on the live data. The process resulted in 541 duplicate editions being deleted from the system and as far as I can tell all is well. We now have 13,086 editions in the system and 13,014 of these do not have an associated book work. We only have 75 book works in the system.
The next step is to assign book works to editions and add in book genres. In order to do this I created a further spreadsheet containing the editions with columns for book work, authors and three columns which can be used to record up to three genres. I also sent Matt and Katie a further spreadsheet containing the details of the 75 existing book works in our system. It’s going to be rather complicated to fill in the spreadsheet as there’s a lot going on and it took me quite a while to figure out a workflow for filling it in. Hopefully with that in place filling it in should be straightforward, if time-consuming.
I also ran some queries, did some checks and generated some spreadsheets for the Wigtown data for Gerry McKeever. With these data related issues out of the way I then returned to developing the front-end. Whilst working on an issue relating to ordering the results by date I noticed that we have quite a lot of borrowing records in the system that have no dates. There are almost 12,000 that don’t have a ‘borrowed year’. There’s possibly a good reason for it, but of these 2,376 have a borrowed day and a borrowed month but no year, which seems more strange. I emailed Katie and Matt about this and they’re going to investigate.
I managed to finish work on the ‘Year borrowed’ bar chart this week. Without providing a year filter the bar chart shows the distribution of borrowing records divided into decades, for example this search for ‘rome’, ordered by date borrowed:
You can then click on one of the decade bars to limit the results to just those in the chosen decade, for example clicking on the ‘1780’ bar:
This then displays a bar chart showing a breakdown of borrowing records per year within the selected decade. You are given the option of clearing the year filter to return to the full view and you can also click on an individual year bar to limit the results to just that year, for example limiting to the year 1788:
When you reach this level no bar chart is displayed as year is the unit that’s filtered and there is only one year selected. But options are given to return to the decade view or clear the year filter. You can of course combine the year filter with any of the other filter options. I guess at year level we could display a similar bar chart for borrowings per month, but this might be too fine-grained and confusing (plus would be a lot more work as everything is currently set up to work with year only). It’s something to consider, though.
I did spot a problem with the bar chart: I realised that when you searched for an individual year or a range within an individual year the results were still showing the options to view the decade and clear the year filter, both of which then gave errors. This has now been sorted – no year filter options should be shown when the main search is only for a year.
For the remainder of the week I began working on the advanced search. As specified in the requirements document, currently the advanced search page features two tabs, one for a ‘simple’ advanced search and one for an ‘advanced’ advanced search. So far I’ve just been working on the forms, which in turn has necessitated making some changes to the API (to bring back a simple list of all libraries and to enable an entire list of registers to be returned). The forms allow you to select / deselect libraries and select / deselect all. In the ‘Simple’ tab there are also textboxes for entering date of borrowing, author forename and surname, year of birth / death and book title, plus a placeholder for genre. The requirements document stated that date of borrowing would have boxes for entering years and days and a drop-down list for selecting month, with two sets to be used for range dates. I’ve decided that since the quick search already allows dates to be entered directly as text that it would make sense to just follow the same method for the advanced search.
Author dates as currently specified are going to be a bit messy for BC dates, where people need to enter a negative value. This is messy because a dash is used for date ranges so we may end up with something like ‘-1000–200’ (that’s two dashes in the middle). I’m not sure what we can do about this, though. I guess having different boxes for ‘from’ and ‘to’ for ranged dates would avoid the issue. For the ‘advanced’ advanced search lists of selectable registers will appear depending on the libraries that are selected. This is what I’m still in the middle of working on.
If I have the time I would like to create a new theme for the website that will look pretty similar but will use the Bootstrap front-end toolkit (https://getbootstrap.com/). The current WordPress theme doesn’t use this which means creating complex layouts is more difficult and messy. I created a Bootstrap based WordPress theme for the Anglo-Norman Dictionary (e.g. this search form: https://anglo-norman.net/textbase-search/) but I’ll just have to see how much time I have as I think it’s better to get the essentials in place first. But what it means is in the meantime things like the search form layout will possibly not be finalised (but will be functional).
In addition to the above I fixed an issue with the Thesaurus of Old English data for Jane Roberts and I completed setting up an initial WordPress site for the VARICS project. I also did a little work for the Dictionaries of the Scots language, fixing a broken link from entries to DOST abbreviations, replying to an email from Rhona about a cookie policy for the website and investigating an issue with text in italics in quotations not being found when a ‘quotations only’ advanced search is performed.
It turns out that the code I’d written to generate the data for the quotations was only set to pick up the direct contents of <q> and to ignore the contents of any child elements such as <i>. This is not the case with the full text and ‘exclude quotations’ data. I identified the issue and updated the code, running a test entry through it to test that the italicised text in quotes is now getting indexed properly. It may well be that there was a reason why the code was set up in this way, though, as Ann mentioned that there are other tags within quotes whose content should be ignored. I’ll need further input from the team before I do anything further about this.
Week Beginning 7th November 2022
I participated in an event about Digital Humanities in the College of Arts that Luca had organised on Monday, at which I discussed the Books and Borrowing project. It was a good event and I hope there will be more like it in future. Luca also discussed a couple of his projects and mentioned that the new Curious Travellers project is using Transkribus (https://readcoop.eu/transkribus/) which is an OCR / text recognition tool for both printed text and handwriting that I’ve been interested in for a while but haven’t yet needed to use for a project. I will be very interested to hear how Curious Travellers gets on with the tool in future. Luca also mentioned a tool called Voyant (https://voyant-tools.org/) that I’d never heard of before that allows you to upload a text and then access many analysis and visualisation tools. It looks like it has a lot of potential and I’ll need to investigate it more thoroughly in future.
Also this week I had to prepare for and participate a candidate shortlisting session for a new systems developer post in the College of Arts and Luca and I had a further meeting with Liz Broe of College of Arts admin about security issues relating to the servers and websites we host. We need to improve the chain of communication from Central IT Services to people like me and Luca so that security issues that are identified can be addressed speedily. As of yet we’ve still not heard anything further from IT Services so I have no idea what these security issues are, whether they actually relate to any websites I’m in charge of and whether these issues relate to the code or the underlying server infrastructure. Hopefully we’ll hear more soon.
The above took a fair bit of time out of my week and I spent most of the remainder of the week working on the Books and Borrowing project. One of the project RAs had spotted an issue with a library register page appearing out of sequence so I spent a little time rectifying that. Other than that I continued to develop the front-end, working on the quick search that I had begun last week and by the end of the week I was still very much in the middle of working through the quick search and the presentation of the search results.
I have an initial version of the search working now and I created an index page on the test site I’m working on that features a quick search box. This is just a temporary page for test purposes – eventually the quick search box will appear in the header of every page. The quick search does now work for both dates using the pattern matching I discussed last week and for all other fields that the quick search needs to cover. For example, you can now view all of the borrowing records with a borrowed date between February 1790 and September 1792 (1790/02-1792/09) which returns 3426 borrowing records. Results are paginated with 100 records per page and options to navigate between pages appear at the top and bottom of each results page.
The search results currently display the complete borrowing record for each result, which is the same layout as you find for borrowing records on a page. The only difference is additional information about the library, register and page the borrowing record appears on can be found at the top of the record. These appear as links and if you press on the page link this will open the page centred on the selected borrowing record. For date searches the borrowing date for each record is highlighted in yellow, as you can see in the screenshot below:
The non-date search also works, but is currently a bit too slow. For example a search for all borrowing records that mention ‘Xenophon’ takes a few seconds to load, which is too long. Currently non-date quick searches do a very simple find and replace to highlight the matched text in all relevant fields. This currently makes the matched text upper case, but I don’t intend to leave it like this. You can also search for things like the ESTC too.
However, there are several things I’m not especially happy about:
- The speed issue: the current approach is just too slow
- Ordering the results: currently there are no ordering options because the non-date quick search performs five different queries that return borrowing IDs and these are then just bundled together. To work out the ordering (such as by date borrowed, by borrower name) many more fields in addition to borrowing ID would need to be returned, potentially for thousands of records and this is going to be too slow with the current data structure
- The search results themselves are a bit overwhelming for users, as you can see from the above screenshot. There is so much data it’s a bit hard to figure out what you’re interested in and I will need input from the project team as to what we should do about this. Should we have a more compact view of results? If so what data should be displayed? The difficulty is if we omit a field that is the only field that includes the user’s search term it’s potentially going to be very confusing
- This wasn’t mentioned in the requirements document I wrote for the front-end, but perhaps we should provide more options for filtering the search results. I’m thinking of facetted searching like you get in online stores: You see the search results and then there are checkboxes that allow you to narrow down the results. For example, we could have checkboxes containing all occupations in the results allowing the user to select one or more. Or we have checkboxes for ‘place of publication’ allowing the user to select ‘London’, or everywhere except ‘London’.
- Also not mentioned, but perhaps we should add some visualisations to the search results too. For example, a bar graph showing the distribution of all borrowing records in the search results over time, or another showing occupations or gender of the borrowings in the search results etc. I feel that we need some sort of summary information as the results themselves are just too detailed to easily get an overall picture of.
I came across the Universal Short Title Catalogue website this week (e.g. https://www.ustc.ac.uk/explore?q=xenophon) it does a lot of the things I’d like to implement (graphs, facetted search results) and it does it all very speedily with a pleasing interface and I think we could learn a lot from this.
Whilst thinking about the speed issues I began experimenting with Apache Solr (https://solr.apache.org/) which is a free search platform that is much faster than a traditional relational database and provides options for facetted searching. We use Solr for the advanced search on the DSL website so I’ve had a bit of experience with it. Next week I’m going to continue to investigate whether we might be better off using it, or whether creating cached tables in our database might be simpler and work just as well for our data. But if we are potentially going to use Solr then we would need to install it on a server at Stirling. Stirling’s IT people might be ok with this (they did allow us to set up a IIIF server for our images, after all) but we’d need to check. I should have a better idea as to whether Solr is what we need by the end of next week, all being well.
Also this week I spent some time working on the Speech Star project. I updated the database to highlight key segments in the ‘target’ field which had been highlighted in the original spreadsheet version of the data by surrounding the segment with bar characters. I’d suggested this as when exporting data from Excel to a CSV file all Excel formatting such as bold text is lost, but unfortunately I hadn’t realised that there may be more than one highlighted segment in the ‘target’ field. This made figuring out how to split the field and apply a CSS style to the necessary characters a little trickier but I got there in the end. After adding in the new extraction code I reprocessed the data, and currently the key segment appears in bold red text, as you can see in the following screenshot:
I also spent some time adding text to several of the ancillary pages of the site, such as the homepage and the ‘about’ page and restructured the menus, grouping the four database pages together under one menu item.
Also this week I tweaked the help text that appears alongside the advanced search on the DSL website and fixed an error with the data of the Thesaurus of Old English website that Jane Roberts had accidentally introduced.
Week Beginning 26th September 2022
I spent most of my time this week getting back into the development of the front-end for the Books and Borrowing project. It’s been a long time since I was able to work on this due to commitments to other projects and also due to there being a lot more for me to do than I was expecting regarding processing images and generating associated data in the project’s content management system over the summer. However, I have been able to get back into the development of the front-end this week and managed to make some pretty good progress. The first thing I did was to make some changes to the ‘libraries’ page based on feedback I received ages ago from the project’s Co-I Matt Sangster. The map of libraries used clustering to group libraries that are close together when the map is zoomed out, but Matt didn’t like this. I therefore removed the clusters and turned the library locations back into regular individual markers. However, it is now rather difficult to distinguish the markers for a number of libraries. For example, the markers for Glasgow and the Hunterian libraries (back when the University was still on the High Street) are on top of each other and you have to zoom in a very long way before you can even tell there are two markers there.
I also updated the tabular view of libraries. Previously the library name was a button that when clicked on opened the library’s page. Now the name is text and there are two buttons underneath. The first one opens the library page while the second pans and zooms the map to the selected library, whilst also scrolling the page to the top of the map. This uses Leaflet’s ‘flyTo’ function which works pretty well, although the map tiles don’t quite load in fast enough for the automatic ‘zoom out, pan and zoom in’ to proceed as smoothly as it ought to.
After that I moved onto the library page, which previously just displayed the map and the library name. I updated the tabs for the various sections to display the number of registers, books and borrowers that are associated with the library. The Introduction page also now features the information recorded about the library that has been entered into the CMS. This includes location information, dates, links to the library etc. Beneath the summary info there is the map, and beneath this is a bar chart showing the number of borrowings per year at the library. Beneath the bar chart you can find the longer textual fields about the library such as descriptions and sources. Here’s a screenshot of the page for St Andrews:
I also worked on the ‘Registers’ tab, which now displays a tabular list of the selected library’s registers, and I also ensured that when you select one of the tabs other than ‘Introduction’ the page automatically scrolls down to the top of the tabs to avoid the need to manually scroll past the header image (but we still may make this narrower eventually). The tabular list of registers can be ordered by any of the columns and includes data on the number of pages, borrowers, books and borrowing records featured in each.
When you open a register the information about it is displayed (e.g. descriptions, dates, stats about the number of books etc referenced in the register) and large thumbnails of each page together with page numbers and the number of records on each page are displayed. The thumbnails are rather large and I could make them smaller, but doing so would mean that all the pages end up looking the same – beige rectangles. The thumbnails are generated on the fly by the IIIF server and the first time a register is loaded it can take a while for the thumbnails to load in. However, generated thumbnails are then cached on the server so subsequent page loads are a lot quicker. Here’s a screenshot of a register page for St Andrews:
One thing I also did was write a script to add in a new ‘pageorder’ field to the ‘page’ database table. I then wrote a script that generated the page order for every page in every register in the system. This picks out the page that has no preceding page and iterates through pages based on the ‘next page’ ID. Previously pages in lists were ordered by their auto-incrementing ID, but this meant that if new pages needed to be inserted for a register they ended up stuck at the end of the list, even though the ‘next’ and ‘previous’ links worked successfully. This new ‘pageorder’ field ensures lists of pages are displayed in the proper order. I’ve updated the CMS to ensure this new field is used when viewing a register, although I haven’t as of yet updated the CMS to regenerate the ‘pageorder’ for a register if new pages are added out of sequence. For now if this happens I’ll need to manually run my script again to update things.
Anyway, back to the front-end: The new ‘pageorder’ is used in the list of pages mentioned above so the thumbnails get displaying in the correct order. I may add pagination to this page, as all of the thumbnails are currently on one page and it can take a while to load, although these days people seem to prefer having long pages rather than having data split over multiple pages.
The final section I worked on was the page for viewing an actual page of the register, and this is still very much in progress. You can open a register page by pressing on its thumbnail and currently you can navigate through the register using the ‘next’ and ‘previous’ buttons or return to the list of pages. I still need to add in a ‘jump to page’ feature here too. As discussed in the requirements document, there will be three views of the page: Text, Image and Text and Image side-by-side. Currently I have implemented the image view only. Pressing on the ‘Image view’ tab opens a zoomable / pannable interface through which the image of the register page can be viewed. You can also make this interface full screen by pressing on the button in the top right. Also, if you’re viewing the image and you use the ‘next’ and ‘previous’ navigation links you will stay on the ‘image’ tab when other pages load. Here’s a screenshot of the ‘image view’ of the page:
Also this week I wrote a three-page requirements document for the redevelopment of the front-ends for the various place-names projects I’ve created using the system originally developed for the Berwickshire place-names project which launched back in 2018. The requirements document proposes some major changes to the front-end, moving to an interface that operates almost entirely within the map and enabling users to search and browse all data from within the map view rather than having to navigate to other pages. I sent the document off to Thomas Clancy, for whom I’m currently developing the systems for two place-names projects (Ayr and Iona) and I’ll just need to wait to hear back from him before I take things further.
I also responded to a query from Marc Alexander about the number of categories in the Thesaurus of Old English, investigated a couple of server issues that were affecting the Glasgow Medical Humanities site, removed all existing place-name elements from the Iona place-names CMS so that the team can start afresh and responded to a query from Eleanor Lawson about the filenames of video files on the Seeing Speech site. I also made some further tweaks to the Speak For Yersel resource ahead of its launch next week. This included adding survey numbers to the survey page and updating the navigation links and writing a script that purges a user and all related data from the system. I ran this to remove all of my test data from the system. If we do need to delete a user in future (either because their data is clearly spam or a malicious attempt to skew the results, or because a user has asked us to remove their data) I can run this script again. I also ran through every single activity on the site to check everything was working correctly. The only thing I noticed is that I hadn’t updated the script to remove the flags for completed surveys when a user logs out, meaning after logging out and creating a new user the ticks for completed surveys were still displaying. I fixed this.
I also fixed a few issues with the Burns mini-site about Kozeluch, including updating the table sort options which had stopped working correctly when I added a new column to the table last week and fixing some typos with the introductory text. I also had a chat with the editor of the Anglo-Norman Dictionary about future developments and responded to a query from Ann Ferguson about the DSL bibliographies. Next week I will continue with the B&B developments.
















