Week Beginning 21st October 2024

I was back at work on Monday this week after being on holiday the week before.  Whilst I was away the external hosting company where many of our sites are now hosted performed a major upgrade to the default PHP version our sites run on, from 7.X to 8.X.  I hadn’t been given an advanced warning of this significant update and unfortunately when it took effect many of our websites stopped working, and the first I was aware of this was when people started sending me frantic emails.  Thankfully I was checking my emails whilst I was away and I spent a bit of my holiday manually reverting the sites back to 7.X, a temporary measure that got the sites up and running again.

On my return to work I spent all of Monday going through every site in order to upgrade any parts of the code that were causing problems with PHP 8.  I had to go through 24 sites, but thankfully in most cases it was the same issue that was causing problems each time, and once identified it was relatively easy to fix.  By the end of the day all of the sites were working properly with PHP 8.X, which was a relief.

The Anglo-Norman Dictionary people had also alerted me to an issue with the dictionary management system whilst I was away too.  When editing entries any language tags were being lost, not from the XML itself, but the entries were no longer being picked up by the language search facility.  I spent a bit of time investigating what was going on here.  When a new or edited entry is published search terms such as language tags are extracted from the entry’s XML.  When I created the new language tag system a few months ago I first applied it to the batch processing scripts that handle an entire batch of entries and once this was tested and working I then I copied the code into the script that handles the publication of entries via the holding area in the DMS.

In each script the extraction of language tags only applies to ‘main’ entries rather than xref entries and there’s a check to ensure the XML file in question is the correct type.  In the batch processing scripts this check uses a variable named $entrytype but in the DMS publication script it’s $pub[“entrytype”] and unfortunately I overlooked this.  The result was that the DMS script was assuming (when processing language tags) that every entry wasn’t a ‘main’ entry and therefore the language extraction section of the script didn’t execute and no language tags were extracted.

It turned out that many entries had been edited since I’d added in the language search, meaning all of these had stopped being findable in the language search.  Thankfully I still had the script I’d originally written to extract language tags from all active entry XML files and populate the language search system, so I was able to re-run this and restore the data.

Also for the AND people this week I had a Teams meeting to discuss the new funding application we’re putting together, and also engaged in several email conversations about this.  I can’t say much more about the application at this stage, but it will involve adding lots of nice visualisations to the site (amongst other things).  I had to spend some time this week researching and considering the technical aspects of the new proposal in addition to the time spent in the meeting and writing emails.

I also spent some further time this week working on the new place-names proposal to the BA with Thomas Clancy and gave some advice to Gavin Miller about a new AHRC proposal he’s putting together.  I also made some minor updates to the Iona place-names map and fixed a bug in the interface that was causing the map legend to appear multiple times if a user clicked the ‘reset map’ several times in quick succession.  When this situation arose, the new legend ended up loading in before the old one was fully removed, causing multiple legends to be displayed.  It took a little time to investigate and my solution is something of a hack but it works:  when you press the reset map button any subsequent clicks in the following second are now ignored.  This gives the map time to sort itself out without getting itself in a tangle.

Whilst investigating this issue I also spotted some issues with the ‘browse placenames’ feature.  When resetting the map the contents of the browse pane were not being reset, meaning the list of browse options would not necessarily match the heading shown.  Also after resetting the map the drop-down list to change the visible browse options was no longer triggering.  I managed to fix both these issues in the Iona map and then roll the changes out to the Ayr map too.

I also discovered this week that LiDAR map layers for some areas of Scotland are publicly available for use in map-based resources.  See this page on the NLS site: https://maps.nls.uk/guides/lidar and specifically the links to the available phases listed here: https://maps.nls.uk/guides/lidar/#re-use.

Unfortunately Iona has not yet been covered, but Phase 3 covers much of the Berwickshire area and KCB while Phase 4 covers much of Ayrshire and I contacted the place-names people to see whether we should think about incorporating the LiDAR data into some of the online resources.  This would need to be handled differently to the other map layers and I think would benefit from being an additional layer that users could turn on or off, sitting atop the base map a user has chosen and including an opacity slider.  This would enable the user to easily compare the LiDAR data to features shown on whichever map they’ve selected.  I reckon it has interesting research potential and the place-names people agreed.  I’ll just need to find the time to add this in.

Also this week I was in communication with the Speak For Yersel people as the new survey areas are going to launch in the next week or so.  We agreed on the steps that would need to be taken (from a technical point of view mainly just updating the live Scotland survey to link through to the other areas) and I also added in some social media links and icons for two of the three new surveys.

Finally, I heard back from Deven Parker this week about her Playbills project and the experiments that have been undertaken by someone in Computing Science using ChatGPT to extract data from the images.  This is all looking very promising.  I can’t say too much more about it at this stage, but I spent some time looking through the sample output and giving some feedback and suggestions to Deven.

Week Beginning 7th October 2024

This was a four-day week for me, as I’d taken Friday off.  I’m also off for all of next week.  I spent a lot of my time this week working on new proposals, or responses to proposals that had been submitted.  I spent much of my time on Monday investigating new ways to visualise the Anglo-Norman Dictionary data for a new proposal that will hopefully lead to the completion of the dictionary.  I also had a lengthy email discussion with the editor Geert about lemmatisation and stemming and how these might be employed in a new tool that the website would offer.  I can’t really say too much more about it for now, but it would be quite an exciting development.

On Tuesday I met with Matthew Creasy to discuss a new proposal to publish the correspondence of Mallarmé.  We had a very productive meeting and I demonstrated the Burns Correspondence resource I’m currently working on, amongst other things.  On Tuesday afternoon I met with Thomas Clancy and Simon Taylor to discuss a new proposal to expand and enhance the Scottish place-names resources.  Again, I can’t say too much about it, but it was a great meeting and I spent a fair amount of time the following day writing some text that will hopefully be incorporated into the proposal.  I also continued to engage with Matt Sangster regarding the feedback to another AHRC proposal.

I did find a bit of time this week to continue with the development of the Burns Correspondence resource.  I managed to solve the issue with filters being forgotten once the timeline loads and the also issue with pressing on a location marker when in the timeline loading the full correspondence for the location rather than the filtered correspondence.  I’ve also now linked to the biography popups from the timeline entries and have ensured that when a red highlighted marker in the timeline map is pressed on it doesn’t deselect (as happens on regular maps) but instead it directly loads the correspondence pane.  Previously in order to reach the correspondence pane from the timeline for a highlighted marker you needed to press the marker twice – once to deselect it and then again to select it, which triggered the load.

I also added in the option to change the base map from the historical one to a satellite map, a fairly unadorned relief map and a modern OS map.  These options are now available on the ‘Home’ tab of the map.  The screenshot below shows the satellite map centred on Mossgiel Farm.  The ‘Home’ tab is now split into two columns and I’ve added some further placeholder text to the left column (taken from the main website):

I also returned to the ‘top tens’ in the facts and figures pages of the Books and Borrowing project to tackle the issue of the top tens not being entirely accurate when multiple (but not all) libraries are selected.  This was caused by the used of cached data: Each library’s ‘top ten’ lists are cached and when multiple (but not all) libraries are selected overall top ten lists are generated by combining these individual ones.  The problem is any borrowing totals that are not in the top ten for a library are not factored into the calculations so (for example) if a book is the eleventh most popular borrowing at one library and the most popular at another then the cound for the former is not included.

I’ve now managed to update the system so that when multiple libraries are selected in the ‘Facts’ page the top ten lists take into consideration all borrowings rather than just those found in the individual top ten lists for each library.  The screenshots below show the top ten book works at Chambers and Advocates.  The following shows ‘Blackwoods’ with a total of 196:

If you look at the ‘results’ screenshot below you’ll see that there are 203 results, with 196 at Chambers and only 7 at Advocates:

As 7 places Blackwoods way outside the top ten for Advocates it wasn’t getting picked up for the amalgamated top ten  The ‘after’ shot shows an accurate figure of 203 for Blackwoods:

I had to make some pretty major changes to get this working.  The cached data for each library now needs to hold the complete list of books (works, editions, holdings) and borrowers together with their counts rather than just the top tens.  This greatly increases the size of the cache files, but thankfully it’s not had too big an impact on the time it takes to load the facts pages.

As the data now includes everything rather than the top tens I also needed to update the way the facts for individual libraries were displayed, as previously the data was simply pulled in and displayed in full, as ‘full’ was already limited to ten items.  Thankfully this is all now sorted (I hope)  and no more than ten items should ever be displayed.

I also responded to a query from Ann at the DSL about bibkiographical notes, and that’s all for this week and until I get back from my holiday.

Week Beginning 30th September 2024

I continued to work on the interactive map of Burns correspondence this week, focussing on the timeline feature.  This feature will take the data and pin it to the map one correspondence at a time, building up a picture of connections over time.  I managed to complete an initial version this week and while there are still some issues with it, it’s mostly working.  I had originally intended that the timeline would automatically animate through each correspondence after a user presses ‘play’, with the option to pause things also offered.  However, as I began working on the feature I came to the conclusion that this would be rather tedious to use: if you’re interested in a particular item you’ll likely not have enough time to consider it before the view whooshes off to the next item, and if you’re not interested in it you have to wait for the next item to trigger.  Instead, I’ve decided to make users manually navigate through the items themselves, which gives users much more control.  I could still potentially investigate adding in the fully automated approach, but I would personally find it a bit annoying to use.

For now the situation is you define the map data (either all of it or limiting the data using the filter option) then you press on the ‘Timeline’ menu item (a good example for testing is selecting ‘Historians’ in the filter options as this gives a manageable 15 correspondence).  When the timeline first loads it clears the map and loads the first item chronologically, centring the map on the ‘from’ location.  After 5 seconds the map automatically flies to the ‘to’ location.  The location markers are highlighted in red, as is the dotted joining line.  The following screenshot demonstrates the timeline with the earliest correspondence in the category ‘Historians’ selected:

The timeline pane features next and previous buttons (previous is greyed out in the screenshot above as we’re looking at the earliest correspondence).  A number is assigned to the item and the total number of items is displayed to give people and idea of where they are (e.g. Item 1 of 15).  The correspondence data is displayed in a box below this, with information about when the item was sent, the sender and recipient and their locations.  For now there is no link to open the bibliography but I’ll add this in.  This is another reason for not automating the runthrough as if we did then opening the bibliography pop-up would potentially mess things up, or it may be difficult to press on a link before it disappears.

Pressing on the ‘Next’ button loads the next item, flying the map to the new ‘from’ location and then flying onto the new ‘to’ location after 5 seconds.  The new markers and line are highlighted in red and the previous markers and line remain on the map, but are turned grey.  Note that for now if you press the ‘Previous’ button the markers and lines for subsequent items continue to be visible on the map (in grey).  I’m not sure whether this is an issue or not, really.  It would be somewhat complicated to remove the items entirely.  It can be done, but it would take some time and I’m not really sure it’s worth it, but I’ll see what people think.

There are some issues that I still need to deal with.  If you press on a marker when in the timeline view this currently loads the correspondence pane but it displays all of the data for the location.  Plus if you go to the filters your current selection is not remembered.  There are also issues with the ‘fly to’ in that often the map tiles don’t load quick enough so until the flying stops you might not see the map.  Also when the distance from the previous ‘to’ location to the current ‘from’ location is quite far the map can begin to fly to the current ‘to’ location before you get much of a glimpse of the ‘from’ location.  Also all of the zooming and panning can get a bit much.  It’s possible a simple jump to the new location might make for a nicer and less queasy experience.

Also this week I spend some time refreshing the cached data for the Books and Borrowing project, as the final researcher has completed work on the data.  It takes some time to generate all of the various caches, but the new data was fully in place by Tuesday.  Matt had also noticed some discrepancies with the statistics in the facts and figures page (specifically the ‘top ten’ lists) and the actual data returned in the search results and I spent some further time investigating this.

The top ten counts were still using the database rather than the Solr index so I updated this, but while this had minor implications for the numbers listed it didn’t bring them into line with the search results.  It turns out that the issue was caused by data truncation.  Top ten summaries are stored for each library and to create an overall top ten these individual lists are then combined.  But if a book does not appear in the top ten for a library (for example if it is number 11) then it’s not found when calculating the overall total so its borrowings are missed and the overall total is therefore incorrect.  I’m sure we’d discussed this issue previously and I thought I’d done something to fix it, but apparently not.  However, I’ve now updated the ‘all libraries’ cache so that it generates and stores top tens directly from across the whole dataset rather than by combining the caches for each library and the default top tens on the facts and figures page should now accurately reflect the number of search results returned. One thing I spotted is that if a user selects multiple libraries (but not all libraries) on the facts and figures page the results are still an amalgamation of each library’s individual top ten lists so there will likely be discrepancies.  We agreed that this would also need addressing so I’m going to look into this next week, all being well.

Also this week I met with Matt to discuss a response to reviewer feedback to an AHRC proposal I’m involved with.  I can’t really say too much more about it, but I needed to spend some time before the meeting reading through the materials and then formulating a response to issues that were relevant to me.  I also met with Thomas Clancy and Simon Taylor to discuss a possible BA application.  We’re going to meet in person next week to take this forward.

I also made a few tweaks to the Place-names of Iona map interface and attended the English Language and Linguistics research seminar on Thursday, which was a very interesting discussion of visualising ‘lone wolf’ terrorists in the British press.