ACM HyperText Site Quality Audit — 2026-09-18
Full review of 1,010 published archive documents, confirming screenshot-only imports, paragraph reconstruction defects, and metadata/navigation inconsistencies.

This is a read-only quality review of the complete published ACM HyperText archive on the Hypermedia Network as it stood on 18 September 2026.

Executive summary

The review mapped all 38 conference-year directories and inspected all 1,010 published child documents. Every published paper URL was fetched successfully. The central problems are real and systematic:

    11 HT’21 papers are page-image imports, not structured papers. Their bodies contain 90 complete-page screenshots and almost no paper text. All 11 source PDFs contain native extractable text—537,950 characters in total—so screenshot publication was unnecessary.

    Paragraph reconstruction is unreliable in a material subset. The broad detector raised 72 severe alerts. PDF-based validation of 16 stratified alerts confirmed at least one true prose-boundary defect in 11 and systemic prose fragmentation in 6. The defects include PDF lines emitted as paragraphs, live sentences split at columns/pages, retained line-wrap hyphens, isolated page numbers, and permission/front-matter inserted into prose order.

    Metadata and navigation contain 46 manually verified findings: 2 high, 42 medium, and 2 low. The root says 1,635 papers; visible year totals sum to 1,695; the published directory has 1,010 children. These may be different populations, but the site does not define them.

    No malformed internal HM paths were found among 5,033 scanned internal links.

The 72 detector alerts are a review queue, not a claim that 72 papers are broken. Short blocks can be legitimate headings, captions, equations, lists, references, and table cells.

P0 — replace page screenshots with structured text

Required fix: reimport title, authors, abstract, headings, paragraphs, lists, tables, captions, references, and real figures as semantic blocks. Preserve the PDF as a downloadable source if desired, but do not use page images as the article body.

P1 — paragraph and reading-order repair

Highest-priority six-document queue

Confirmed defect patterns


    Line-per-paragraph serialization: PDF display lines become separate semantic paragraphs.

    Column and page splits: one sentence is broken into two blocks at a column or page transition.

    Dehyphenation failure: fragments such as node- / based, doc- / continuation, and informa- / tion remain split.

    Reading-order failure: permission notices, authors, page numbers, figures, or captions interrupt the body and the sentence resumes later.

    Over-segmentation of furniture: isolated words and page numbers are represented as normal paragraphs.

P1 — reconcile archive counts and HT’26 navigation

The root and year pages mix at least three populations without labeling them: proceedings records, mirrored/full-text papers, and published Seed child documents.

    Site root claims 1,635 papers; its visible year rows sum to 1,695 items; year directories contain 1,010 published children.

    The root says 148 papers are mirrored, while the only visible per-year mirrored annotations sum to 147 (37 + 58 + 52).

    HT’26 says 61 identified records, 55 imported, and 6 blocked; its directory has 60 published children, but only 55 child links appear on the landing page.

    34 year pages make item-count claims that differ from directory child counts. A difference can be legitimate, but every page should label the two populations and expose missing, metadata-only, blocked, and full-text states.

Required fix: create one machine-readable per-paper manifest with conference year, DOI, record state, Seed document URL, full-text/mirror status, license, and blocker reason. Derive root and year totals from it.

P2 — metadata and presentation corrections

Manually verified corrections:

    ECHT’90: “Knowledfe” → “Knowledge”; item titles lack DOI links and standard author presentation.

    HT’00: “San Antionio” → “San Antonio”.

    HT’03: start year 20003; use 26–30 August 2003.

    HT’05: end year 92005; use 6–9 September 2005.

    HT’15: dates incorrectly use 2014; use 1–4 September 2015.

    HT’22: end year incorrectly uses 2023; use 28 June–1 July 2022.

    HT’26: author names are plain text despite the root promise that every paper links to its authors.

    Paper-title navigation changes silently by era: DOI links before 2020, Seed links after 2020. Use a consistent dual-link convention.

    The public root embeds internal-looking Augmentation Process notes. Move operational notes out of the reader-facing landing page.

Remediation order

    Rebuild the 11 screenshot papers from native text.

    Repair and republish the six score-71 papers; then review the 13 score-63 papers.

    Add the per-paper manifest and regenerate all counts, statuses, and year landing lists.

    Correct dates, typos, DOI links, author links, and the public-root process embed.

    Run automated gates on every future import and on every repaired document.

Import-quality gate

A paper should fail publication when any of these conditions holds:

    complete-page images are used while the PDF exposes substantial native text;

    median prose block length resembles display lines and short-block ratios are high;

    consecutive blocks match a hyphenated or lowercase syntactic continuation;

    permission text, page numbers, author matter, figures, or captions interrupt a sentence;

    title/author/abstract/section/reference coverage is incomplete;

    source-to-publication figure, table, equation, and paragraph counts have not been reviewed;

    exact published read-back has not been compared with the prepared artifact.

Method and limits

The audit combined a complete directory crawl, full rendered-document block metrics, source-PDF inspection with PyMuPDF, and manual coordinate-aware comparison of a stratified severe sample. It did not treat the detector score as ground truth. The observed validation precision is useful for triage but is not an unbiased estimate of archive-wide prevalence.

The archive-wide evidence and per-document metrics were retained as structured audit artifacts for remediation and reruns.

Do you like what you are reading? Subscribe to receive updates.

Unsubscribe anytime