The Two Machines
Appraisal, access, and the ceiling on the computational archive
A MetaArchivist essay
Artificial intelligence promises to transform archival discovery. But every AI system built atop an archive inherits a decision almost no one discussing AI ever examines: most records were destroyed before the archive ever saw them. The future of computational access is therefore bounded not by search technology but by appraisal — and appraisal itself is now becoming computational.
That is the whole argument. The rest of this essay explains why it is true, why it has been hiding in plain sight, and why it reframes the question that will define the next twenty years of digital archives.
Throughout this essay, I use computational to mean systems that do more than retrieve records: they derive, infer, or reason over relationships among records — through structured queries, knowledge graphs, entity extraction, or machine learning.
Four conceptions of access
Traditional archival theory distinguishes between physical and intellectual control. What follows is a different framework: not how archivists control records, but how researchers gain access to them.
Access at the U.S. National Archives has evolved through four distinct ideas of what “getting to the record” means, each layered on top of the last rather than replacing it.
The first is physical access: finding aids and reading rooms, the researcher traveling to the records and consulting a paper surrogate to locate a box. The second is descriptive access: those surrogates migrated online — the Archival Research Catalog, retired in 2013; the Online Public Access prototype that succeeded it; and the fully rebuilt National Archives Catalog that launched in November 2022, keyed on the National Archives Identifier.[1] The third is digital access: not a description of the record but the preservation-grade digital object itself, delivered through the Electronic Records Archives repository. The fourth, now arriving, is computational access: knowledge graphs, entity extraction, relationship discovery, machine reasoning — access not to a record but to the relationships among records.
That progression does real work. It places ERA in its proper historical position — not the endpoint of digital archiving but the bridge between a descriptive archive and a computational one. And it explains the shape of the present moment, in which every discovery tool, including AI, is being asked to work from descriptive metadata alone.
One feature of that present moment will matter throughout what follows. The systems that make up the archive share identifiers, not understanding. Each preserves its own view of the archive, but no layer models relationships across them. The implications of that architectural choice become visible only when we ask what AI is actually reasoning over.
But the four-stage story has a floor it never mentions — and, as we will see, the obvious rejoinder to this essay is that the floor could itself be pulled up into the graph. Hold that objection; it is the right one, and I will return to it. First the floor.
Appraisal: the universe over which access operates
Before any of the four kinds of access can operate, a prior decision has been made: whether the record is permanent. That decision is not another access layer. It determines the universe over which every later access technology operates. In the U.S. system, it is encoded in a records schedule — a disposition authority, carrying the identifier prefixed DAA when processed through ERA’s scheduling function.[2]
Here the personal history of the system matters, because it explains why the system looks the way it does. ERA did not reinvent appraisal for the digital age. It computerized the analog appraisal model the National Archives had refined over decades, preserving the assumptions about permanence, provenance, and disposition that had governed federal appraisal for decades while translating them into machine-executable workflows. The architecture inherited the philosophy intact. The DAA became a machine-readable object instead of a paper schedule; transfer increasingly became the transfer of electronic records rather than paper files — though agencies continued, and in some cases still continue, to transfer physical media such as hard drives; disposition became workflow rather than correspondence. But the underlying intellectual framework remained Jenkinson, Schellenberg, and the Federal Records Act — not a new theory built for computational archives. It was not an oversight that appraisal sits beneath computational access. It was a design inheritance from the paper era.
ERA did not reinvent appraisal for the digital age.
The precision is worth getting right, because it sharpens the point. A disposition authority is not a mark of preservation. A single DAA schedule routinely covers both permanent and temporary items; the authority simply states, item by item, what shall happen to each. For the overwhelming majority of federal records, the authority orders destruction, on a defined clock, carried out by the agency or a Federal Records Center. Only the sliver of items dispositioned permanent is ever transferred to the National Archives.
So NARA ordinarily never has to destroy temporary federal records, because they are destroyed before transfer; the temporary records — the vast majority of everything the government creates — simply never arrive. Reappraisal remains possible, however, and NARA may authorize later destruction of accessioned records under limited circumstances. NARA’s own long-standing estimate is that permanent records amount to roughly one to three percent of what the federal government produces; one frequently cited internal analysis put the figure near 1.23 percent.[3] Whatever the exact number, the archive is a small, deliberate residue of the whole, and it is selected upstream, before ERA’s repository ever touches it.
Now return to the four stages and notice what they share. Physical, descriptive, digital, and computational access all range over that same residue. Not one concerns the records scheduled temporary and destroyed. “Access” has always meant access to what appraisal kept — four increasingly sophisticated ways of reaching into a set whose boundary was drawn somewhere else, by someone else, for reasons that had nothing to do with discovery.
The machine that split
ERA assumed the existing appraisal process was correct and built software to execute it more efficiently. Although ERA’s original vision included improved public access, the architecture that was ultimately delivered centered on transfer, preservation, authenticity, scheduling, and repository management. Discovery remained largely downstream of those functions — treated as a consumer of preserved records rather than a co-equal design objective. That is why discovery was never central to what shipped.
There is a deeper way to see it. The analog archive had one machine: the archivist. The same profession decided what entered the archives and helped researchers find it, and there was an implicit unity between those functions because they lived in the same people. ERA dissolved that unity institutionally, into separate automated pipelines — an appraisal machine (scheduling, transfer, disposition), a preservation machine (the repository), an access machine (the Catalog), and a structured-query machine (AAD). The historical unity of archival judgment was split into independent systems, each optimizing something different. That split is the precondition for everything that follows.
The analog archive had one machine: the archivist.
Why this is a ceiling, not a gap
It would be comfortable to treat the appraisal filter as a gap that better engineering might close. It is not a gap. It is a ceiling.
The value framework of appraisal is not the value framework of discovery. Records are appraised series by series, often decades before anyone imagines the questions a future researcher — or a future model — might ask, and they are judged on evidential and informational value under a provenance-based theory of the archive. That theory has real virtues: it respects the context of creation, keeps like with like, treats arrangement itself as meaning. But it is a theory about custody and evidence, not relationships and inquiry. A computational layer built atop the permanent residue inherits this appraisal ontology completely, and cannot see it. The relationship that would have connected a destroyed temporary series to a surviving permanent one is not underrepresented in the graph. It is absent — structurally and permanently — because one of its endpoints was never admitted.
You can discover every relationship you like among the permanent records. You cannot discover the ones that ran through what was thrown away. Machine reasoning over the archive reasons over a set whose edges were defined by a judgment it has no access to and could not question if it did.
The proof it was always possible — and the reason it stayed boxed in
Two artifacts from NARA’s own history make this concrete, and they are easy to confuse because they share three letters in different orders.
The first is AAD, the Access to Archival Databases system, launched in 2003. AAD lets the public run fielded queries — by name, date, place, organization — across hundreds of structured database files drawn from dozens of federal agencies. It is, in every meaningful sense, computational access. And here is the fact that should stop you: AAD was the first publicly accessible application developed under the auspices of the Electronic Records Archives program.[4] Its engineering had in fact begun earlier — NARA contracted the work to SAIC in the summer of 1999 — but the ERA program claimed it as its inaugural public deliverable. The first public demonstration of ERA’s future was not preservation but computation. Computational access did not arrive as the fourth and latest stage. It appeared first, in 2003, and was then subordinated as ERA proper became a preservation system and the descriptive Catalog became the front door.
Why did it stay boxed in? Because AAD worked precisely where the records were already structured data — and even there, only over datasets individually processed and packaged expressly for AAD, not across NARA’s electronic holdings as a whole. The descriptive layer — the surrogates standing in for billions of unstructured records — never inherited that queryability. AAD demonstrated computational access over a carefully prepared subset of structured datasets; it did not, and could not, generalize to the archive at large.
The second artifact is the DAA disposition authority itself, and it explains the architecture’s priorities. ERA’s organizing spine is the schedule. Its first institutional job is appraisal and disposition — the machine-actionable decision about permanence and destruction. A system whose primary key is a disposition authority is optimized for custody and accountability, not discovery. The trusted digital repository was the goal; the computational discovery platform was not. AAD is the capability; DAA is the reason it never became the center of gravity.
A federation that shares no mind
Look at the systems actually in operation, and the ceiling stops being abstract. NARA runs not one system or two but a loosely coupled federation.
The Catalog, successor to ARC and OPA, holds the descriptions and exposes them through an online user interface and an open read-write API — but the API serves descriptive metadata and crowd contributions, tags and transcriptions, not relationships.[5] AAD still holds the structured-data query portal, separate from the Catalog. ERA, now rebuilt as ERA 2.0 on a modular, microservices architecture in AWS GovCloud, holds scheduling and the preservation of permanent electronic records, with distinct workflows for federal, presidential, legislative, judicial, digitized, and donated materials.[6] ARCIS, the Archives and Records Centers Information System, holds the physical custody and logistics of the Federal Records Centers — the pre-digital layer, still running.[7] And a ring of engagement surfaces — DocsTeach for education, History Hub for crowdsourced reference, Founders Online for curated documentary editing — sit atop the descriptions.
What these systems do not share is a knowledge layer. They share identifiers, not understanding. Original order is preserved within each fonds; almost nothing models relationships across them. This is the stage-four ceiling made concrete: not a missing feature but a missing mind. There is no place in the architecture where the structured records in AAD, the preservation objects in ERA, the custody records in ARCIS, and the descriptions in the Catalog are related to one another as a graph.
Two machines, computerizing in opposite directions
Here the argument turns. Look at what NARA’s own artificial-intelligence program is doing. Its public inventory of AI use cases describes automated tagging of roughly two million digital records, natural-language search over internal holdings, and generative tools to summarize and draft against operational records.[8] Each applies machine learning to the descriptive layer, and each produces more description.
AI is not replacing archival description. It is industrializing it. Automated tagging, transcription, summarization, and semantic enrichment all manufacture more description — better metadata fed to the same descriptive front door. Intelligence is being bolted onto access.
The disposition side is a different kind of machine, and the difference is the crux. Appraisal itself — the judgment that a body of records is permanent or temporary — remains human work. NARA’s appraisal archivists negotiate schedules with agency records officers, series by series, under the same evidential-value framework the profession has used for decades; where machine learning is being piloted to help categorize records, it recommends, and a person still decides. No model decides what history is worth keeping. What automates is not the judgment but its execution. The schedule becomes machine-actionable: role-based approaches like Capstone let an agency dispose of records by the position of the account holder rather than by reading them, approved schedules run on a clock, and destruction becomes workflow.[9] NARA’s machine-implementable General Records Schedule makes the mechanism concrete. Archivists still decide what is permanent and what is temporary; those approved decisions are then decomposed into standardized data elements — a disposition, a retention period, an event trigger — and published as a software-agnostic CSV that recordkeeping platforms such as Microsoft Purview can ingest to automate disposition.[10] NARA even refines the format through a monthly, NARA-led Microsoft 365 User Group; one 2023 update added a field “in response to suggestions at a recent Microsoft 365 User Group meeting.”[11] The appraisal judgment remained human; only its execution became portable and machine-actionable. The human decides the rule once; machines then enforce it at a scale no reading room ever imposed.
That is what the two-machines image is actually about. It is not that AI now appraises. It is that the two ends are being mechanized in opposite ways. The access machine is being made intelligent — it discovers relationships among the records that survived. The disposition machine is being made automatic — it enforces, at volume, the human decisions about which records never survive. And the judgment that once unified both functions in a single archivist now sits between them, authoring disposition rules on one side and description standards on the other, while automation carries out each in isolation. The archive’s silences — its negative space — are still decided by people; they are increasingly operationalized by machine, executed against records that never enter the visible corpus.
So there are two machines. One, on the access side, grows smarter at discovering relationships among the records that survived. The other, on the disposition side, grows more automatic at enforcing which records never survive. Both are computerizing at once, and they do not speak to each other.
This is the sentence the whole essay has been walking toward: the future of access is capped by the past of appraisal, and both ends are now computerizing independently — the machine that finds and the machine that forgets, never in the same conversation.
The successor question, restated
The obvious question is whether ERA’s successor should expose its internal knowledge graph directly, rather than forcing every future discovery interface — including AI — to work from descriptive metadata alone. The answer is almost certainly yes. But it is not the deep question, because it accepts the residue as given. It optimizes access to what appraisal already kept.
Now the objection I asked you to hold. Could the appraisal decisions themselves become nodes in the graph — disposition authorities modeled as first-class data: explicit, queryable objects in their own right rather than hidden implementation details, so that the archive at least records the shape of its own exclusions? Technically, yes, and that is exactly the point.
That is the choice that will define the next twenty years of digital archives. Not whether ERA’s successor exposes a graph, but whether that graph has any awareness of the far larger graph appraisal declined to keep. ERA, on this reading, is neither the endpoint of digital archives nor merely the bridge to computational ones. It is the seam between two machines that have not yet been introduced.
Federal records management has traditionally treated evidence of authorized destruction as a new records problem rather than as archival context to preserve — the very instinct a computational archive will have to revisit. Every archive preserves evidence of what was kept. The computational archive will also need to preserve evidence of what was intentionally forgotten. Until then, machine reasoning will mistake archival residue for historical reality.
Notes
On the lineage from the Archival Research Catalog (retired 2013) through the Online Public Access prototype to the National Archives Catalog (launched November 2022) and the role of the National Archives Identifier: National Archives, “The National Archives Catalog — About,” https://www.archives.gov/research/catalog/about, and “National Archives Catalog Enhancements,” https://www.archives.gov/research/catalog/ngc-preview.
On disposition authority numbers, records schedules, and ERA’s scheduling function: National Archives, “Disposition Authority Number” (Lifecycle Data Requirements Guide element), https://www.archives.gov/research/catalog/lcdrg/elements/disposition-authority-number; “Records Control Schedules,” https://www.archives.gov/records-mgmt/rcs; and the NARA Records Management glossary, https://www.archives.gov/files/records-mgmt/rm-glossary-of-terms.pdf.
On the share of federal records judged permanent (commonly stated as one to three percent; one internal analysis near 1.23 percent): National Archives, “Appraisal Policy of the National Archives,” https://www.archives.gov/records-mgmt/scheduling/appraisal, and “The Percentage of Permanent Records in the National Archives: A 1985 Article Revisited,” The Text Message blog, https://text-message.blogs.archives.gov/2020/04/14/the-percentage-of-permanent-records-in-the-national-archives-a-1985-article-revisited/. Estimates vary by agency and period but consistently represent only a small fraction of records created.
That AAD (launched April 2003) was “the first publicly accessible application developed under the auspices of” the ERA program is stated verbatim in both National Archives, “AAD: A New Tool to Search NARA Databases,” Prologue (Spring 2003), https://www.archives.gov/publications/prologue/2003/spring/spotlight-aad.html, and press release nr03-34, “Thousands Search National Archives New Electronic Database” (April 8, 2003), https://www.archives.gov/press/press-releases/2003/nr03-34. The Prologue article — written by David R. Kepley, AAD’s project manager since its inception — notes that development began in the summer of 1999 under a contract with Science Applications International Corporation (SAIC), and that AAD addressed “just access to a specific type of electronic record—databases,” while NARA was separately “concentrating efforts on the development of the ERA system.” So the attribution is one of program sponsorship, not authorship by a distinct ERA engineering office. See also the AAD Privacy Impact Assessment, https://www.archives.gov/files/privacy/privacy-impact-assessments/aad.pdf.
On the open, read-write Catalog API: National Archives, “API for the National Archives Catalog,” https://www.archives.gov/research/catalog/help/api, and the Catalog-API repository, https://github.com/usnationalarchives/Catalog-API.
On ERA 2.0’s modular, microservices, AWS GovCloud architecture and multiple record-type workflows: National Archives Open Government Plan, “Electronic Records Archives,” https://usnationalarchives.github.io/opengovplan/erecordsarchives/, and industry summary, https://www.tagovcloud.com/2023/09/naras-electronic-records-archives-2-update-primer/.
On ARCIS and the Federal Records Centers: National Archives, “About Archives and Records Center Information System (ARCIS),” https://www.archives.gov/frc/arcis/about.
On NARA’s AI use cases — automated tagging of roughly two million records, natural-language search, and generative productivity tools: National Archives, “Inventory of NARA Artificial Intelligence (AI) Use Cases,” https://www.archives.gov/ai, and “New Strategic Framework Emphasizes … Responsible Use of Artificial Intelligence,” https://www.archives.gov/news/articles/new-strategic-framework-artificial-intelligence.
On appraisal as human judgment and the rule-based (rather than machine-learning) automation of disposition: National Archives, “Records Scheduling and Appraisal,” https://www.archives.gov/records-mgmt/sch-appraisal, and “Scheduling Records,” https://www.archives.gov/records-mgmt/scheduling/sch-records. On Capstone role-based scheduling: National Archives, “White Paper on The Capstone Approach and Capstone GRS,” https://www.archives.gov/files/records-mgmt/email-management/final-capstone-white-paper.pdf. On NARA’s exploratory review of machine learning for records work: National Archives, “Cognitive Technologies White Paper: Records Management Implications,” https://www.archives.gov/files/records-mgmt/policy/nara-cognitive-technologies-whitepaper.pdf.
On the machine-implementable GRS as a data model — each disposition authority decomposed into a disposition (permanent/temporary), a retention period, and a standardized event trigger, published as a software-agnostic CSV to “at least partially automate disposition”: National Archives, “Machine-Implementable GRS,” https://www.archives.gov/records-mgmt/grs/machine-implementable-grs, and “FAQs About the GRS Machine-Implementable Format” (record layout), https://www.archives.gov/records-mgmt/grs/machine-implementable-faq. The FAQ notes that some event types can be automated while “others may require human action.” Commercial platforms consume such rules through CSV file-plan import: Microsoft, “Records management for documents and emails in Microsoft 365,” https://learn.microsoft.com/en-us/purview/records-management.
On NARA’s refinement of the format through a standing Microsoft forum: National Archives, AC 41.2023, “Updated Machine-Implementable General Records Schedule (GRS) File” (July 26, 2023), https://www.archives.gov/records-mgmt/memos/ac-41-2023 — “we made this update in response to suggestions at a recent Microsoft 365 User Group meeting” — and “Communities of Interest,” https://www.archives.gov/records-mgmt/policy/coi, which describes the NARA-led Microsoft 365 User Group that “meets monthly to discuss topics around the records management implications of the Microsoft 365 software platform.”
If this essay changed how you think about archives, consider becoming a paid subscriber. Paid subscriptions help support the research, archival digging, standards work, and long-form writing behind essays like this. They make it possible to keep asking the questions that sit beneath the technology, not just describing the technology itself.



I come from a corporate environment, and one of the biggest problem we have serving users is spending time looking for things that are not in the archive because they were never supposed to be there, but the user, who is not privy to those disposition decisions, never knew that.
In many corporate environments, that kind of decisioning is never written down. I can count on all my fingers and toes the business people I have spoken to who lack knowledge management systems and even simple governance guides, let alone business rules documentation. I once became the only operational employee of a startup that had never even printed an employee handbook or kept an archive of email to refer to the decisions the executive team had made using that channel. So I can tell you that outside of academic and GLAM+ environments, it is extremely hard to find documentation of disposition decisions, even when there is a DAM Librarian in-house or an Operations Director who is conscious and careful about the documentation of the decisions they make, and tries to write documentation for every new workflow or business rule.
Things can change fast in corporations, even in the archive. You'd think things like metadata schemas, taxonomies and ontologies would remain static...but they simply do not. One new executive vice-president can make it necessary to change everything. That is going to happen. And it needs to be traced and every relationship change tracked in a knowledge layer that is separate from assets but easily discoverable and searchable, because the questions will come up.
"Every archive preserves evidence of what was kept. The computational archive will also need to preserve evidence of what was intentionally forgotten." For the evidence of what was intentionally forgotten, what evidence is needed beyond the disposition authority and possibly the appraisal justification for that authority? Archives rarely have sufficient resources to preserve and provide access to their permanent holdings. What additional burdens are you suggesting they take on?
Why does the computational archive "need to preserve evidence of what was intentionally forgotten"? What knowledge do you envision the end user gaining from such evidence?
History has always been written based on whatever records have survived until the time the history is written, and not all of those records will be in the custody of archives. The historian rarely accesses all of the available evidence.