How to Resolve EPUBCheck HTM_004 and HTM_011 Entity Errors in EPUB 3?

How to Resolve EPUBCheck HTM_004 and HTM_011 Entity Errors in EPUB 3?

EPUB 3.3 defines XHTML Content Documents strictly as XML serializations conforming to the W3C Extensible Markup Language (XML) 1.0 (Fifth Edition) specification. Under XML 1.0 §4.6, only five character entity references are predefined: &amp;, &lt;, &gt;, &quot;, and &apos;. All other named entity references (e.g., &nbsp;, &mdash;, &hellip;) are illegal unless explicitly bound to a system or public identifier via a Document Type Definition (<!DOCTYPE>).

Commercial conversion pipelines regularly emit legacy HTML5 named entities into EPUB containers. When evaluated against the schema rules of EPUBCheck, this triggers a fatal parsing error:

ERROR(HTM-004): book.epub/OEBPS/ch01.xhtml (24,112): The entity "nbsp" was referenced, but not declared.

To resolve HTM-004, developers frequently inject traditional XHTML 1.1 or transitional DTD headers:

<!DOCTYPE html PUBLIC "-//W3C//DTD XHTML 1.1//EN" "http://www.w3.org/TR/xhtml11/DTD/xhtml11.dtd">

This introduces a secondary failure mode under EPUB 3.3 §4.2.1:

ERROR(HTM-011): book.epub/OEBPS/ch01.xhtml (2,85): An external entity declaration was found in the DOCTYPE declaration. External entities are forbidden in EPUB 3.

EPUB 3.3 explicitly prohibits external DTD subsets and external parameter entities to eliminate network round-trip overhead during reading system ingestion and to close the security surface for XML External Entity (XXE) injection attacks. Internal DTD subsets are technically valid in generic XML 1.0, but EPUB 3 rendering engines either reject internal entity expansions or exhibit undefined behavior.

The standard resolution requires two constraints:

  • Retaining only the minimal HTML5 doctype: <!DOCTYPE html>.
  • Normalizing all non-predefined named character entities to their hex/decimal Unicode Numeric Character Reference (NCR) equivalents (e.g., &#160; or &#xA0;) or direct UTF-8 encoded code points.

Pipeline Architecture

The remediation pipeline operates as a deterministic pre-packaging filter within continuous integration runs:

  • Unpack / Stream Ingestion: Read uncompressed XHTML source streams directly from disk or prior pipeline stages (e.g., Pandoc output).
  • DOCTYPE Stripping/Replacement: Strip legacy PUBLIC and SYSTEM DTD identifier blocks using regex token matchers, resetting the header strictly to <!DOCTYPE html>.
  • Lexical Entity Interception: Tokenize byte sequences matching &[A-Za-z0-9]+;.
  • Reserved Exclusion: Filter out core XML entities: &amp;, &lt;, &gt;, &quot;, &apos;.
  • Unicode Mapping: Resolve remaining tokens via standard HTML5 character entity lookup tables, replacing them in-place with hexadecimal numeric character references (&#x...;).
  • OPF Manifest Verification: Inspect the .opf manifest to verify that XHTML spine items carry media-type="application/xhtml+xml" without invalid fallback declarations or missing XML namespace prefixes (xmlns="[http://www.w3.org/1999/xhtml](http://www.w3.org/1999/xhtml)").
  • Preflight Assertion: Trigger headless EPUBCheck CLI validation; halt the pipeline with non-zero exit codes on any schema or parser regressions.
+------------------+     +-----------------------+     +--------------------------+
| Raw Pandoc/HTML  | --> | Entity & DOCTYPE      | --> | EPUB Container Packager  |
| XHTML Generators |     | Sanitizer (Python/CLI)|     | (ZIP -0 / Standard Store)|
+------------------+     +-----------------------+     +--------------------------+
                                                                    |
                                                                    v
                                                       +--------------------------+
                                                       | EPUBCheck 5.1 CLI Audit  |
                                                       | Assertion: Exit Code 0   |
                                                       +--------------------------+

Executable Implementation

The following production script, epub_entity_sanitizer.py, traverses an unpacked EPUB staging directory, normalizes all XHTML content files in place, verifies the OPF manifest, and repacks the output.

#!/usr/bin/env python3
"""
EPUB 3 Entity Sanitizer and Preflight Pipeline
Resolves EPUBCheck HTM_004 and HTM_011 violations.
"""

from __future__ import annotations

import html.entities
import os
import re
import sys
from pathlib import Path
import xml.etree.ElementTree as ET

# Predefined XML entities that must not be converted to numeric references
XML_PREDEFINED = {"amp", "lt", "gt", "quot", "apos"}

# Regex matching named character entities: &name;
ENTITY_PATTERN = re.compile(r"&([a-zA-Z0-9]+);")

# Regex targeting any DOCTYPE declaration with external subsets
DOCTYPE_PATTERN = re.compile(
    r"<!DOCTYPE\s+html[^>]*>",
    re.IGNORECASE | re.DOTALL,
)


def entity_replacer(match: re.Match[str]) -> str:
    """Resolve named HTML entity into a hexadecimal numeric character reference."""
    entity_name = match.group(1)
    if entity_name in XML_PREDEFINED:
        return match.group(0)

    codepoint = html.entities.name2codepoint.get(entity_name)
    if codepoint is not None:
        return f"&#x{codepoint:04X};"

    # Fail hard if an unknown/unmapped entity enters the pipeline
    raise ValueError(f"Unresolvable entity token encountered: &{entity_name};")


def sanitize_xhtml_content(content: str) -> str:
    """Normalize DOCTYPE declarations and rewrite entities to hex NCRs."""
    # Enforce pure EPUB 3 DOCTYPE without external subsets
    sanitized = DOCTYPE_PATTERN.sub("<!DOCTYPE html>", content)
    # Transmute named entities to numeric representations
    sanitized = ENTITY_PATTERN.sub(entity_replacer, sanitized)
    return sanitized


def sanitize_directory(oebps_dir: Path) -> int:
    """Scan and process all .xhtml and .html content documents in place."""
    processed_count = 0
    extensions = {".xhtml", ".html", ".htm"}

    for path in oebps_dir.rglob("*"):
        if path.suffix.lower() in extensions and path.is_file():
            raw_text = path.read_text(encoding="utf-8")
            transformed_text = sanitize_xhtml_content(raw_text)

            if raw_text != transformed_text:
                path.write_text(transformed_text, encoding="utf-8", newline="\n")
                processed_count += 1

    return processed_count


def verify_opf_manifest(opf_path: Path) -> None:
    """Verify MIME types and namespace integrity in the package manifest."""
    if not opf_path.is_file():
        raise FileNotFoundError(f"OPF manifest not found at: {opf_path}")

    # Use explicit namespaces to prevent ElementTree parse defaults from breaking
    namespaces = {"opf": "http://www.idpf.org/2007/opf"}
    tree = ET.parse(opf_path)
    root = tree.getroot()

    manifest = root.find("opf:manifest", namespaces)
    if manifest is None:
        raise ValueError("Malformed OPF: <manifest> node missing.")

    for item in manifest.findall("opf:item", namespaces):
        href = item.get("href", "")
        media_type = item.get("media-type", "")

        if href.endswith((".xhtml", ".html")) and media_type != "application/xhtml+xml":
            raise ValueError(
                f"Invalid media-type '{media_type}' for XHTML document: {href}. "
                "Must be 'application/xhtml+xml'."
            )


def main() -> None:
    if len(sys.argv) < 2:
        sys.stderr.write("Usage: epub_entity_sanitizer.py <unpacked_epub_dir>\n")
        sys.exit(1)

    root_dir = Path(sys.argv[1]).resolve()
    if not root_dir.is_dir():
        sys.stderr.write(f"Error: Directory does not exist: {root_dir}\n")
        sys.exit(1)

    print(f"[*] Processing EPUB sources in: {root_dir}")

    # 1. Sanitize text documents
    modified_files = sanitize_directory(root_dir)
    print(f"[*] Rewrote entities across {modified_files} document(s).")

    # 2. Inspect OPF manifests
    opf_files = list(root_dir.rglob("*.opf"))
    if not opf_files:
        sys.stderr.write("[-] Build Failure: No .opf manifest discovered.\n")
        sys.exit(1)

    for opf in opf_files:
        verify_opf_manifest(opf)
        print(f"[*] Verified OPF manifest integrity: {opf.name}")

    print("[+] Sanitization and manifest inspection completed successfully.")


if __name__ == "__main__":
    main()

Automated Validation

After transforming the source files, package the EPUB container using strict OCF (Open Container Format) physical standards enforcing that the mimetype file is stored uncompressed as the first entry in the archive and assert compliance using EPUBCheck:

#!/usr/bin/env bash
set -euo pipefail

STAGING_DIR="./dist/book_unpacked"
OUTPUT_EPUB="./dist/book_validated.epub"

# 1. Run python in-place entity and manifest sanitizer
python3 ./epub_entity_sanitizer.py "${STAGING_DIR}"

# 2. Re-pack container enforcing OCF specifications
cd "${STAGING_DIR}"
rm -f "${OUTPUT_EPUB}"

# Store mimetype uncompressed (-0) and without extra attributes (-X)
zip -0 -X "${OUTPUT_EPUB}" mimetype

# Deflate remaining payload
zip -9 -r -X "${OUTPUT_EPUB}" * -x mimetype -x "*.DS_Store"
cd - > /dev/null

# 3. EPUBCheck CLI Preflight Assertion
echo "[*] Launching EPUBCheck 5.1 assertion..."
epubcheck "${OUTPUT_EPUB}" \
  --mode exp \
  --failonwarnings \
  --quiet

echo "[+] PASS: Zero entity or schema violations detected."

Production Gotchas

  • CDATA Escaping Traps: Automated regex replacers without boundary awareness will mutate string literals within embedded <script> or <style> CDATA blocks. If raw JavaScript containing logical AND operations (if (a && b)) was improperly serialized, converting bare & tokens or malformed strings inside CDATA causes XML parser termination. Ensure styling and scripting are separated into external assets (.css, .js).
  • Lossy Decimal NCR Mapping in MathML: MathML elements (<math>) embedded in XHTML documents require explicit entity conversion. Replacing MathML-specific named references (such as &InvisibleTimes; or &ApplyFunction;) using a legacy HTML4 table drops characters or maps them into U+FFFD (Replacement Character). The entity transformer must utilize the complete HTML5 / MathML 3 entity map containing all 2,125 named character references.
  • MIME Type Mismatches in the Package Document: Converting files to XHTML standards without matching the OPF manifest entry will pass character entity checks but fail during reading system spine parsing. If an item is declared as text/html instead of application/xhtml+xml, EPUBCheck will raise OPF-012 (non-standard media type without fallback).

You may also like

See All Posts →