Newspapers represent one of the richest available primary sources for historical research — capturing not just major events, but births and deaths, iconic sports victories, political debates, advertisements, and letters to the editor by residents from all walks of life. For historians, genealogists, and other curious minds, newspapers provide an invaluable record for understanding our past.
The Boston Public Library (BPL) has spent decades preserving and providing access to these resources. BPL holds tens of millions of pages of historical newspapers on microfilm, roughly three million of which have been digitized and made searchable online through DigitalCommonwealth.org.
Now, BPL and the Institutional Data Initiative (IDI) at the Harvard Law School Library are releasing a new dataset drawn from more than 1.47 million scanned newspaper pages — a window into 135 years of the public record, from 1795 to 1930. This release represents the next stage in making these collections even more useful for research and discovery.
The Challenge with Existing Digitization
As part of the digitization process for newspapers, software converts page images into searchable text — a process called optical character recognition, or OCR. The problem is that historical newspapers are notoriously difficult for this software to handle. Dense columns, decorative typefaces, and the ravages of age mean the extracted text is often garbled or incomplete. As researchers from IDI note in the technical report accompanying the dataset, “for many newspaper collections, the extracted text is of such low quality that it serves primarily as a loose and unreliable keyword search used to locate the original page scans from which a human can then read, rather than as an accurate representation of the underlying source.”
A second challenge is that newspapers are typically digitized and accessed at the issue or page level, but patrons need to be able to easily locate relevant articles, rather than having to laboriously scan through an entire page image to find the desired content. Patrons are also typically looking for certain types of articles, such as business listings or marriage records.
A New Approach
With these issues in mind, BPL librarians and IDI data scientists designed the approach together, starting from a question BPL brought to the project: What would it mean to understand a digitized newspaper page at the level of its individual parts, not just as a whole image?
Using machine learning, the project isolated more than 83 million individual articles and segments within those pages, then ran text extraction on each one separately — a more targeted approach than processing whole pages at once. Across the full dataset, that process produced billions of words of extracted text.
The collection spans Massachusetts broadly, with titles from Boston neighborhoods like Charlestown, Dorchester, and Roxbury alongside papers from Worcester, Springfield, Salem, New Bedford, and beyond. The dataset also reflects the diversity of the region's communities: while most content is in English, it also includes newspapers in Yiddish, German, Swedish, and French.
For each of those 83 million items, the dataset includes not just higher-accuracy text, but also what kind of content it is — such as news articles, advertisements, birth notices, literary works, and illustrations — plus the names of people, places, and organizations mentioned.
Why It Matters
Once the enhanced data has been integrated into BPL’s online newspaper collections, the most immediate benefits will be an improved, more precise search. Article-level discovery means a researcher looking for coverage of a specific event, or a genealogist searching for an ancestor's death notice, no longer has to browse through entire issues to find what they want.
The richer data also opens up new kinds of research: tracking how “viral” stories spread from paper to paper across the region, or tracing how language and opinions shifted over decades in immigrant-community newspapers.
Longer term, the dataset could serve as a foundation for AI tools capable of answering conversational questions — "What was the public reaction to the Great Boston Fire?" or "Show me sports coverage from the 1918 World Series."
The open-source processing pipeline built for this project is available for any library or archive to adapt for its own newspaper collections. Designed from the outset to run on workstation-level hardware, it can help unlock millions of pages of historical newspapers with greater accuracy than costly commercial solutions, at a fraction of the cost.
Explore the Dataset and Collections
The pipeline code, dataset, and full technical documentation are all openly available:
- Dataset: huggingface.co/collections/institutional/institutional-newspapers
- Technical report: "Institutional Newspapers Pipeline: Deriving billions of high quality tokens from historical newspapers"
- Code: github.com/institutional/institutional-newspapers-pipeline
Read IDI's announcement about the dataset and pipeline at institutional.org.
BPL's digitized historical newspaper collections can be browsed at DigitalCommonwealth.org.
OCR Example
Example from The Evening Union (Springfield, Mass.), December 28, 1910:

Traditional OCR:
“| —————; .
| Funeral of Mrs. E. M. Hannan.
The funeral of Mrs. Eva M. Han-
nan was held from the home of her
sister in law, Mrs. Clifford a eyo,
Oscar
yesterday at 2 o'clock, Rev,”
Enhanced OCR:
“Funeral of Mrs. E. M. Hannan.
The funeral of Mrs. Eva M. Hannan was held from the home of her sister in law, Mrs. Clifford Prevost, yesterday at 2 o'clock, Rev. C. Oscar Ford officiating. The body was taken to Hazardville, Conn., by special car, where burial will take place in the family lot.”

