“Just Because We Can Doesn’t Mean We Should”
Open Data, Messy Machines, and a Few Lessons Learned
It was the after-lunch slot at the Archives and Records Association 2025 Conference, Next Generation: Innovation and Imagination in Record Keeping, held in Bristol on August 27, when the Research Data librarians from the University of Bristol bounded onto the stage determined to keep us awake. Their paper title set the tone: “Just because we can doesn’t mean we should: prototyping machine learning tools to monitor and assess research data.” And to be fair, talking about reproducibility and machine learning was not the dull grind you might expect.
The trio introduced themselves with a mix of authority and personality. Dr. Kirsty Merrett, a Research Support Librarian with 26 years at the University of Bristol, has been leading their research data management programme and publishing open and controlled access datasets in the university repository. She is a strong advocate for FAIR principles, especially in the tricky realm of qualitative data. Alongside her was Jade Godsall, once a medievalist, now Assistant Research Support Librarian at Bristol, whose career path has taken her from manuscripts to machine learning, with side interests in digital preservation and accessibility. And rounding out the team was Christopher Warren, another Assistant Research Support Librarian, who spends his days supporting the research data lifecycle and his evenings raising a young family.
Their theme was simple but provocative: just because we can measure reproducibility, should we?
They walked us through the UK Reproducibility Network pilots, eight projects testing whether institutions and solution providers could collaborate on developing prototype machine-learning tools to assess open research practices. The aims sounded tidy, create valid, reliable, ethical indicators for things like data availability, pre-registration, and credit. But as they pointed out, indicators only have meaning if they reflect real practices, not just algorithmic cleverness.
The ethical stakes were spelled out. Too often, tools show improvements in the algorithms, not in the culture. If dashboards tell you openness is “improving,” but it is only because the software has got better at parsing clunky statements, that is a hollow victory. Their spectrum of openness, from “dog-level open” (recalling an old Lycos ad where even a Labrador could surf the web) to the dreaded “dark data,” made the point memorably.
“Controlled statements matter.”
— University of Bristol Research Data team
Christopher then took us into the weeds of methodology. The team pulled together more than 2,600 records and handed them to three providers. What came back was almost comical: one reported 299 usable results, another 1,187, another 2,672. “Slight variation,” he joked. The machines did fine when the text was clean and simple, but once qualitative data or museum archives entered the picture, they stumbled. Entire repositories went unrecognized. In one striking case, 77 percent of datasets that should have been marked “on request from the institution” were instead misclassified as “on request from the author.”
Jade closed with a practical reminder: controlled statements matter. When researchers used a standard template for their data availability statement, machine-learning tools picked it up with 93 percent accuracy. Without the template, accuracy plummeted to 64 percent. Her punchline was clear: if we want tools to work, we need consistency in how we describe and publish data.
The trio wrapped up with a warning. Machine learning can do the easy stuff, but if we lean on it too heavily, too soon, we risk penalizing entire disciplines, especially the humanities, social sciences, and GLAM fields, whose data does not fit neat technical categories. Ethical monitoring still needs human judgment.
I sat there mulling over a familiar thought: this is exactly the kind of messy, cross-disciplinary challenge where standards could help. Consistent templates, agreed vocabularies, and metadata in predictable places are not glamorous, but they are what make interoperability possible. And here, I could not help but think of ISO TC 46, particularly SC 9 (Identification and Description). SC 9 has long been in the business of building the scaffolding that allows information to be discovered, cited, and trusted: persistent identifiers, metadata schemas, interoperability models, and controlled vocabularies. Their work underpins everything from DOIs for research outputs (ISO 26324), ISNIs for creators (ISO 27729), and ISBNs for books (ISO 2108), to emerging identifier frameworks for digital objects and authority data.
But it is not just about identification. SC 4 (Technical interoperability) is the part of TC 46 that looks at how systems actually talk to each other. They maintain the protocols and exchange formats that allow repositories and catalogs to connect. Think of ISO 23950 (Z39.50) for information retrieval, ISO 2709 for bibliographic record exchange, or ISO 2146 for registry services. These are the kinds of technical specifications that let structured metadata travel across systems and be harvested reliably.
When the Bristol team pointed out that inconsistent data statements, uncontrolled language, and patchy metadata placement make automated monitoring unreliable, I could not help but hear a call for SC 4 as well as SC 9. The identifiers and descriptive frameworks need to be there, yes, but so do the technical specifications that ensure those frameworks are implemented consistently across platforms.
The frustrations voiced in this session, repositories slipping under the radar, free-text statements that resist automation, whole disciplines whose data is invisible to machine learning, fall squarely in the space between SC 9 and SC 4. If research data is going to be findable, reusable, and machine-actionable across disciplines, it needs exactly the kind of international agreement on identifiers, metadata placement, and exchange protocols that these subcommittees specialize in.
Because reproducibility is not just about data science or machine learning. It is about governance, structure, and community discipline. And that is where the standards world, and particularly the tandem work of SC 9 on identification and description and SC 4 on interoperability, has a real role to play.


