Arts · Language · Gutenberg

LCSH: 69,027 Headings Anchor Gutenberg's Core

LCSH accounts for 69% of subject headings; PS (American literature) appears 4,684 times — concentration defines the reusable public-domain canon.
Kyle McAuliffe · June 15, 2026 · 4 min

Library of Congress Subject Headings (lcsh) account for 69,027 records in a 100,000-row TidyTuesday extract of Project Gutenberg — 69% of all subject entries. The catalog is not a popularity contest; it is a map of which public-domain books volunteers could digitize and classify at scale.

Subject code PS — American literature — recurs 4,684 times, leading all labels. A short head of repeated classifications anchors the catalog while most subject entities appear once. The pattern is typical of library metadata: a reusable canon of headings, and a long inventory of singleton assignments.

Reading Gutenberg as a concentration map of reusable works is where the data help.

Public-domain books cluster by subject type

Public-domain books cluster by subject type

lcsh dominates with 69,027 records, far ahead of thinner subject-type buckets. The main classification family carries the story; this field does not split into many equal long-tail types. That concentration means most navigational claims about what Gutenberg contains are really claims about how LCSH-style headings organize the digitized shelf.

A small set of subjects anchors the catalog

A small set of subjects anchors the catalog

PS appears 4,684 times — the most recurring subject name in the file. The top dozen account for a visible share of all 100,000 rows even though most subject entities appear only once. American literature codes, fiction labels, and related headings form a reusable core. They are the catalog's gravitational center for teachers, scrapers, and adaptation hunters.

Subject families show the catalog center of gravity

Subject families show the catalog's center of gravity

lcsh is again the largest bucket on the category chart. Subject families show where editorial attention should focus first if the goal is to understand the shelf's center of gravity rather than its exotic edges. The edges still matter for discovery. They do not define the statistical middle of a 100,000-row extract.

Repeated subjects reveal the reusable canon

Repeated subjects reveal the reusable canon

Most subject entities appear only once; a small head recurs repeatedly. That power-law shape is typical of catalog tables: a reusable canon of headings, and a long inventory of singleton classifications. Repeated subjects are the ones most likely to support classroom packs, themed collections, and machine-learning corpora. Frequency is a reuse forecast as much as a shelf description.

Subject labels become the map of the shelf

Subject labels become the map of the shelf

PS and related labels become the map of the shelf when numeric scores are sparse. Frequency leaders reveal franchise depth in literature the way studio logos reveal franchise depth in film. The practical claim for cultural analytics is simple: if you can only afford to study a slice of Gutenberg, the repeated subject head is where coverage of the reusable public-domain canon begins.

What this file cannot tell you

Community-cleaned TidyTuesday snapshots are not live APIs. Missing values, spelling variants, and sampling limits apply. Subject headings are not sales, downloads, or reading-time proof.

Findings describe structural signals about Project Gutenberg subject metadata — not a complete history of world literature, and not a ranking of artistic merit.

What to take away

Gutenberg's subject catalog is concentrated: lcsh dominates subject type, a short head of codes such as PS recurs thousands of times, and most labels appear once.

The citable lesson is about reusable canon. Public-domain literature becomes infrastructure when classification and digitization make the same subjects easy to find again and again.

Data, methods & sources

Data and method

The source is the TidyTuesday Project Gutenberg release from the R for Data Science community. The working file contains 100,000 rows after assembly — subject types, subject labels, and related catalog metadata.

Because many fields are categorical, the analysis leans on counts and repetition rather than a single quality score. Charts export as Plotly JSON with PNG fallbacks. Subject headings are librarian infrastructure, not reader reviews.

Sources

Data Science Learning Community. (2025). TidyTuesday: Project Gutenberg (curated via {lcshr}). https://github.com/rfordatascience/tidytuesday/tree/main/data/2025/2025-06-03

Data Science Learning Community. (2025). TidyTuesday: lcsh_subjects.csv. https://raw.githubusercontent.com/rfordatascience/tidytuesday/main/data/2025/2025-06-03/lcsh_subjects.csv — subject types and labels (lcsh, lcc, and related fields) used for the counts in this report.

Fenner, M. (n.d.). lcshr: Download and Process Public Domain Works from Project Gutenberg. rOpenSci. https://docs.ropensci.org/lcshr/

Project Gutenberg. (n.d.). Project Gutenberg. https://www.gutenberg.org/

Library of Congress. (n.d.). Library of Congress Subject Headings (LCSH). https://www.loc.gov/aba/cataloging/subject/headings/