Caselaw Access Project (Harvard Law School)
The complete digitized collection of US case law, containing over 6.4 million cases spanning 1658 to 2018.
The Caselaw Access Project (CAP) was a landmark digitization initiative by Harvard Law School’s Library Innovation Lab that converted the complete US case law collection — over 40 million pages across 6.4 million cases — into structured, machine-readable data. Spanning from 1658 to 2018, CAP covers federal and state appellate courts. The full dataset is available for bulk download, and a research API provides programmatic access. All data is released under a CC0 waiver.
Data Structure
Each case includes the full text of the opinion, standardized case metadata (parties, court, docket number, date), headnotes, and citation information including parallel citations. Data is provided in JSON and XML formats. The CAP API allows searching by citation, full-text query, jurisdiction, court level, and date range. Bulk downloads are available via Amazon S3.
Research Applications
CAP has been used extensively in empirical legal research, natural language processing, and legal analytics. Researchers have employed the dataset for studies on judicial behavior, citation analysis, precedent dynamics, and legal language evolution. The dataset’s temporal scope enables longitudinal studies of legal doctrine spanning centuries. CAP data is frequently combined with CourtListener, Oyez, and other datasets for comprehensive legal analysis.