The new dataset expands the existing Corpus of Founding Era American English (COFEA) by roughly 20%, bringing the total collection to nearly seven million words. Unlike previous resources focused primarily on formal convention records, this update captures the voices of newspaper editors, pamphleteers, and political advocates who shaped public opinion during the ratification debates. By linking entries directly to the University of Wisconsin Libraries, the platform allows scholars to transition seamlessly between statistical linguistic analysis and original primary-source documents.
BYU Law Expands Founding-Era Archive with 7 Million New Words
Timed for Constitution Day, BYU Law has integrated the Documentary History of the Ratification of the Constitution into its corpus linguistics platform. This addition of 14,000 primary texts offers researchers a massive, searchable dataset to decode the original public meaning of words during the birth of American governance.

David Armond, assistant dean for IT at BYU Law, noted that the project provides judges and attorneys with a clearer picture of how constitutional language functioned in 18th-century discourse. The integration arrives as the legal profession grapples with the role of artificial intelligence in research. While AI tools assisted in the digitization and quality control of these historical texts, the corpus itself remains a transparent, evidence-based alternative to black-box algorithmic analysis. BYU Law’s platform has already influenced federal jurisprudence, most notably in United States v. Escobar-Temal, where Sixth Circuit Judge Amul Thapar utilized COFEA data to interpret 18th-century usage patterns.




Comments (0)
No comments yet. Be the first!