Week Beginning 23rd October 2023

After a delightful holiday last week I was back at work again this week.  This involved spending quite a bit of time catching up with emails and dealing with the ongoing issue of migrating sites from old servers to either our new external supplier or a newer server hosted internally.  I was involved with the migration of the SCOTS Corpus to a new server, with my work including fixing a few PHP errors that were cropping up on the more up to date server.  There were also some issues relating to database connections as the original code (which I didn’t write) uses rather a lot of connections – more than the new server was set to allow.  We had thought we’d fixed the issue but it looks like further investigation will be required.

We also migrated the thesaurus.ac.uk site and the Bilingual Thesaurus of Everyday Life in Medieval England (https://thesaurus.ac.uk/bth/) to a new server, which also required tweaking some of the code.  The new server was caching scripts that generated different output each time they were run (e.g. to generate the random category on the homepage), meaning the category wasn’t random but was constantly stuck on ‘Lard a roast’, which wasn’t very helpful.  Thankfully we managed to unstick the cache.

Also this week I investigated an issue with the advanced search of the Dictionaries of the Scots Language as the full-text search had stopped working.  It turned out that the Solr index that powers this search had entirely disappeared from the server, which is more than a little concerning.  It wasn’t a huge issue to rectify as I had the configuration scripts and the data on my PC, but we’re in the dark as to how the index could have been removed.  It had also been brought to my attention that some of the video files I’d uploaded for the Speech Star project before I went on holiday had also disappeared and I’ve reuploaded them too.  Our IT people are investigating what might have caused these issues and if they are linked, but it is concerning.

I also spent a bit of time looking through the old arts.gla.ac.uk server to try and figure out what needed to be retained from it.  It’s mostly old subject area sites that were long ago superseded by T4, plus old conference sites that are no longer needed.  A few of the other sites I’ve already previously moved to T4 myself (e.g. https://www.gla.ac.uk/schools/critical/aboutus/resources/stella/projects/starn/  and https://www.gla.ac.uk/schools/critical/aboutus/resources/stella/projects/bibliography-of-scottish-literature/).  The only site that I think need to be retained are the STELLA apps that I developed from old teaching resources in around 2015.  I therefore requested a new subdomain be set up to host them and migrated them over.  I’ve also requested we set up external hosting for arts.gla.ac.uk, purely to host redirects from old URLs so we don’t end up with broken links.  The new sites are now available (see https://stella.glasgow.ac.uk/aries/, https://stella.glasgow.ac.uk/grammar/, https://stella.glasgow.ac.uk/eoe/, https://stella.glasgow.ac.uk/metre/ and https://stella.glasgow.ac.uk/readings/) but  the redirects from the old URLs are not yet in place.  I’d really like to spend some time redeveloping all of these old apps (apart from ARIES, which has already been redeveloped).  Maybe next year I’ll find some time.

I also set up a new project website for Rhona Brown in Scottish Literature.  I’ve created a bare-bones website at the moment and I’m awaiting further instruction from her on things like themes, colour schemes, site structure and logos.  I also tweaked the project website I’d set up a couple of weeks ago for Petra Poncarova in Scottish Literature to improve the URLs for the Gaelic version of the homepage and helped a project team member get access to a site I’d set up for Matthew Creasy in English Literature.

On Wednesday morning this week I participated in a networking event for the new Research Professional Staff Network.  The event went well and it was very interesting to find out more about other people involved in research support across the University.

For the remainder of the week I began work on the development of the new ‘map first’ interface for the place-names projects, which I’m developing initially for the Iona project.  Below is a screenshot of how things look so far:

At the moment the interface consists of a narrow bar at the top of the browser window with the site’s icon, title and subtitle using the blue colour of the site’s banner as a background.  You can press on the logo or site title to navigate to the main site.  The rest of the browser is taken up with the map.  On the left is the side menu.  As discussed in the requirements document I previously wrote, it consists of four collapsible sections, with ‘Home’ open by default.  I haven’t had the time to implement the search and browse options yet, but the ‘Display options’ section is operational, as you can see above.  Pressing on the section’s title will open the section and you can access the various options.  You can show or hide the side menu by pressing on the button above it.

For the moment the map displays all data that has been marked as ‘on web’ in the CMS (362 records, I think).  By default these are colour-coded by classification code.  The legend is displayed in the top right, allowing you to turn specific features on or off.  You can also show or hide the legend to free up space.  In the bottom right are zoom options plus a ‘full screen’ button that does what you’d expect.  You can press on a map marker to open up the pop-up.  As of yet there is no link through to the full record and some Gaelic fields may be visible.  These will be removed at some point.

Using the ‘Display options’ in the side menu you can change how the map markers are classified.  We may need to be a little more fine-grained with start date and especially altitude.  Also colours for classification codes are currently arbitrarily assigned but we might want to change this – having blue for ‘field’ seems a bit daft, for example.  You can also change the base map and these options are currently the same as for the other place-name sites.  We still need to figure out if / how we can integrate another map of Iona that we discussed at a meeting before I went on holiday.  There is also an option to turn labels on or off.

That’s as far as I’ve got this week.  There’s still a lot to do but I’ve made pretty good progress.  I’ll hopefully find some time to continue with this next week.  I also discovered that the Leaflet mapping library has a method to set the map view so as to show all markers at the closest zoom possible so I’ll ensure I use this when I develop the search and the browse.  I’m currently already using it when the map is first opened to ensure that all of Iona, Soa in the south-west and Eilean Annraidh in the north-east are always visible, no matter what dimensions your screen / browse window are.

Week Beginning 9th October 2023

I came down with some sort of flu-like illness last Friday evening and was still unwell on Monday and unable to work.  Thankfully I was well enough to work again on Tuesday, although getting through the day was hard work.  I was also off on holiday on Friday this week so only ended up working three days.  I’ll be on holiday all of next week as well as it’s the school half-term and we have a family holiday booked.

I was involved in the migration of the Historical Thesaurus website to a new server for a lot of this week.  This required a lot of testing of the newly migrated site and a significant number of small updates to the code to ensure everything worked properly.  Thankfully by Thursday all was working well and I was able to go on my holiday without worrying about the site.

Also this week I did some further work on the Books and Borrowing project, which included generating several different spreadsheets of book holdings that have no associated borrowing records and discussing the options of creating downloadable bundles of all data associated with each specific library.

I also did some work for the Dictionaries of the Scots Language, including investigating an issue with the new quotations search that is not yet live but is running on our test server.  A phrase search for quotations was now working, but an identical phrase search using the full-text index was working fine.  This was a bit of a strange one as it looks like the new Solr quotation search is not picking up the fact that a phrase search is being run.  I tried running the search directly on the Solr instance I’d set up on my laptop and the same thing was happening: I gave it a phrase surrounded by double quotes but these were being ignored.  An identical search on the fulltext Solr index picked up the presence of quotes and successfully performed a search for the phrase.  The only difference between the two fields is that the fulltext fields was set to ‘text_general’ while the quote search was set to ‘text_en’.  I therefore set up a new version of the quote index with the field set to ‘text_general’ and this solved the problem.  I’m still in the dark as to why, though, and I can’t find any information online about the issue.

I also responded to a request from Craig Lamont in Scottish Literature about a new proposal he’s putting together.  If it gets funded I’ll be involved with the project, making a website, an interactivem map and a timeline.  I also had a conversation with Rhona Brown about the website for her new project, which I’ll set up after I’m back from my holiday.

Week Beginning 2nd October 2023

This was a week of many different projects.  On Monday I completed work on a new project website for Petra Poncarova in Scottish Literature, and it is now publicly accessible (see https://erskine.glasgow.ac.uk/).  I also added a blog page to Ophira Gamliel’s project website, created a page for their first blog post (now available here: https://himuje-malabar.glasgow.ac.uk/reconnecting-the-split-moon/) and updated the site to include a link to the blog in the site menu.  This required shifting a few things around to make room for the new menu item.  I also investigated an issue Luca was having in migrating one of Graeme Cannon’s old websites which was similarly structured to the House of Fraser Archive site and managed to find the section of code that was causing the problem (a flag in a regular expression that has since been deprecated).

On Tuesday I completed my work on the CSV endpoints for the Books and Borrowing project, ensuring all nested arrays are ‘flattened’ when producing the two-dimensional CSV file.  This has been a lengthy and tedious task, but it’s good that it’s done, and it should mean that future researchers will be able to extract and reuse the data in a relatively straightforward manner.

On Wednesday I met Luca and Stevie, two of my fellow College of Arts developers to have a catch-up, which was hugely useful as always.  We’ll hopefully meet up again in the next couple of months.  I also responded to a request from Luca to help get some screenshots ready for print publication.  Screenshots are generally 72DPI but this is too low for print.  I’ve previously got around this using Photoshop by loading the image then going to image -> image size.  In the options you can then untick ‘Resample Image’ and then update the ‘resolution’ to whatever you want.  I’ve never actually printed the resulting images to check any difference, but I’ve never had anyone come back and ask for better versions.  I guess another option would be to take the screenshots on something like an iPad that natively runs at a higher DPI.

Also on Wednesday I spent some time on the DSL, investigating an issue with Google Analytics for Pauline Graham and then investigating a problem with phrase searching and highlighting that Pauline had also noticed on both the live and test sites.  When a phrase was searched for each individual word in the phrase was being highlighted in the entry, and then if you returned to the search results and went back to an entry from there no highlighting worked.  Also some search results were not featuring snippets.  This turned out to be three separate issues that needed to be investigated and fixed:

  1. Separate word highlighting: The default setting in the highlighting library I installed a few months ago highlighted each word in a string.  If there were multiple words separated by spaces then all matching words would be highlighted.  Thankfully the library (https://markjs.io/) has a setting that only matches the entire string and I’ve activated this now.  Now if you perform a search for ‘off or on’ or something and navigate to a result only the exact term will be highlighted.
  2. Losing the highlighting when navigating back to the results and then to an entry: This was a problem with spaces getting encoded between pages.  They were becoming the URL encoded equivalent ‘%B’ or ‘+’ and after that the string no longer matched.  I’ve sorted this.
  3. Lack of snippets: The issue was down to the length of the entry.  In Solr, the snippet generation is a separate process to the search matching.  While the search checks the entire entry the snippet generation by default only looks at the first 51,200 characters.  An entry such as ‘Mak’ is a long entry and if the search term only matches text quite far down the entry a snippet doesn’t get created.  After discovering this I’ve updated the setting so that 100,000 characters are analysed instead and this has fixed the issue.  More information about this can be found at https://stackoverflow.com/questions/52511154/solr-empty-highlight-entry-on-match.

This investigation took some of Thursday as well, after which I moved back to the Books and Borrowing project, for which I spent some time generating data relating to the Royal High School for checking purposes.  I also received some bid documentation for a proposal Gavin Miller is putting together.  Gavin wanted me to read through the documentation and add in some further sections relating to the data.  The data will consist of a directory of projects and resources which will be available to search and browse, plus will be visualised on an interactive map.  I added in some information and hopefully the proposal is a success.

On Friday I made some further updates to the Speech Star websites, adding in some new videos to the Edinburgh MRI Modelled Speech Corpus (https://www.seeingspeech.ac.uk/speechstar/edinburgh-mri-modelled-speech-corpus/) and arranging their layout a bit better.  I also replied to a request from Rhona Brown, who would like a website to be set up for a new project she’s starting work on soon.  I listed a few options we could pursue and I need to wait to hear more from her now.

I also spent quite of bit of time investigating some minor issues Ann Ferguson had spotted with the predictive search on the DSL website, most of which will thankfully be sorted when the new Solr based headword search goes live.

Finally, I had a meeting with the Placenames of Iona project to discuss the development of a new ‘map first’ interface for the data.  I met with Thomas, Sofia and Alasdair and it was really great to actually have an in person meeting with them, having never done so before.  We discussed many aspects of the interface and had some really useful discussions.  I’ll be starting on the development of the front-end in the coming weeks.

Week Beginning 25th September 2023

I had my PDR session on Monday this week, which was all very positive.  There was also one further UCU strike day on Wednesday this week, cutting my working days down to four.  The project I devoted the most of the available time to was Books and Borrowing.  Last week I had begun reworking the API to make it more usable and this week I completed this task, adding in a few endpoints that I’d created but hadn’t added to the documentation.  I then moved onto the task of adding ‘Download data’ links to the front-end.  These links now appear as buttons beside the ‘Cite’ button on any page that displays data, as you can see in the following screenshot:

Pressing on the button loads the API endpoint used to return the data found on the page with ‘CSV’ rather than ‘JSON’ selected as the file type.  This then prompts the file to be downloaded by the browser rather than loading the data into the browser tab.  It took a bit of time to add these links to every required page on the site, but I think I’ve got them all.  However, the CSV downloads still needed quite a lot of work doing to them.  When formatted as JSON any data held in nested arrays are properly transformed and usable, but a CSV is a flat file consisting of columns and rows and the data has a more complicated structure than this.  For example, if we have one row in the CSV file for each borrowing record on a register page the record may have multiple associated borrowers, each with any number of occupations consisting of multiple fields.  The record’s book holding may have any number of book items and may be associated with multiple book editions and there may be multiple authors associated with any level of book record (item, holding, edition and work).  Representing this structure in a simple two-dimensional spreadsheet is very tricky and requires the data to be ‘flattened’.  In order to do so a script needs to work out the maximum number of each variable items a record in the returned data has in order to create the required columns (with heading labels) and to pad out any other records that don’t have the maximum number of items with empty columns so that the columns of all records line up.

So, for example, when looking at borrowers: If borrowing row number 16 out of 20 has a borrower with five occupations then column headings need to be added for five sets of occupation columns and the data for the remaining 19 rows needs to be padded out with empty data to ensure any columns that appear after occupations continue to line up.  As a borrowing may involve multiple borrowers this then becomes even more complicated.

I managed to update the API to ensure nested arrays were flattened for several of the most complicated endpoints, such as a page of records and the search results.  The resulting CSV files can become quite monstrously large, with over 200 columns of data a regular occurrence.  However, with the data properly structured and labelled it should hopefully make it easier for users who are interested in the data to download the CSV and then delete the columns they are not interested in, resulting in a more manageable file.  I still need to complete the ‘flattening’ of CSV data for a few other endpoints, which I hope to tackle next week.

Also this week I had an email discussion with Petra Poncarova, a researcher in Scottish Literature who is beginning a research project and requires a project website.  I’ve arranged for hosting to be set up for this and by the end of the week we had the desired subdomain and WordPress installation.  I spent a bit of time on Friday afternoon getting the structure and plugins in place and next week I’ll work on the interface for the website.

I also made a couple of further updates to the House of Fraser Archive this week.  I’d completed most of the work last week but hadn’t managed to get the search facility working.  After some suggestions from Luca I managed to figure out what the problem was (it turned out to be the date search part of the query that was broken) and the search is now operational.  We even managed to get results highlighting in the records working again, which is something I wasn’t sure we’d be able to do.

The rest of my time was spent making updates to the Speech Star websites (and Seeing Speech).  Eleanor had noticed some errors in the metadata for a couple of the videos in the IPA charts so I fixed these.  There were also some better quality videos to add to the ExtIPA charts and some further updates to the metadata here too.  Also for this project Jane Stuart-Smith contacted me to say that I had been erroneously categorised as ‘Directly Incurred’ rather than ‘Directly Allocated’ when the grant application had been processed, which is now causing some bother.  I may have to create timesheets for my work on the project, but we’ll see what transpires.