Our methodologies for good quality scientific practice are improving every day. For years I attended Tim Williams' post-excavation methods talks when he was a visiting lecturer in UCC (in the early 1990s) and absorbed his whole 'groups and subgroups' methodology and we built it into our professional practice as field archaeologists. Archaeological recording methods combine indexes, registers, record sheets, scaled plans, scaled photos, sampling and timelines. All of these are usually corralled by the simplest of tools, word documents and spreadsheets. We rarely used more than 2-3 heading types in our reports and we kept our spreadsheets simple - flat file databases essentially.
Our training was to objectively present our findings in such a way that other archaeologists could disagree with us and use our own evidence to support their hypotheses. And this system works very well, it is not perfect but it aspires to a structured data approach which is intellectually honest.
Sound principles build good quality data which results in good quality work. This AI/LLM revolution we're experiencing now is a valuable tool which will augment such sound practices. Those simple word docs and excel sheets are ideally structured for AI parsing and analysis. AI methods for working with such data are being published regularly and some such as this recent article outline a very clear AI-facing method for working with our types of reports. The AI parses our reports into the smallest units of verifiable data without making interpretative leaps and then it groups them semantically.
It's applying similar methods to the excavation and post-excavation methodologies developed in the 1980s and 1990s and it is doing it at scale. I took that pdf and in a single pass got an Ai to write two scripts to apply their methodology (I literally said to Gemini - read this pdf and propose a plan for writing scripts based on it) - I then tested the scripts on a single report and it condensed a five page report into 21 'verifiable claims' and then into four interpretative groups.
There's more testing to do but we are almost at the point where I can send a WhatsApp/Telegram message to an OpenClaw bot to trawl our work servers - extract every word doc tagged #BronzeAge (for example) and run this Booeshaghi, Luebbert and Pachter inspired prompt over them. I won't have to sit down and pull everything together; from what I have seen of the bot so far it is like a dog with a bone - it will go off and hunt these files and when it is done I will get a cheerful WhatsApp message telling me there is a new dataset based on our own good quality structured data ready for examination. Based on the rate of development this will be working next week, not next year.
This focus on good principles leading to good practice and high quality data works just as much for community heritage projects. Combine the deep local expertise of local archaeologists/historians/genealogists with professional practitioners as we have experienced since 2010 and add the technological multiplier of AI and the prospects are good for high quality collaborations. Is machine readble heritage good? As we say in Irish "gan dabht"!!!
(h/t to our Jack for the Booeshaghi et al reference)
