Catalog TraceAcademic data, with evidenceOpen original catalog

Independent Python + Playwright sample

From a course catalog
to data you can trust.

Discover the public course response, normalize its records and inspect every result against its source.

Inspect the results

Loading records

Extracted courses. Select a row to inspect its evidence.
CourseCreditsChecks

A small pipeline.
A complete chain of evidence.

Built for the practical work of catalog discovery, extraction, transformation and debugging.

  1. Discover

    Playwright opens the public Coursedog catalog and observes its course response and actual course links.

  2. Normalize

    Python cleans HTML, preserves credit ranges and creates stable course identifiers. Missing values stay explicit.

  3. Validate

    Invalid required fields and conflicting duplicates stop publication. Each exported record retains its source URL.

  4. Deliver

    CSV, JSON, JSONL, JSON-LD, a read-only REST API and an offline snapshot. No model API is required.

Run it. Review it. Reproduce it.

The downloadable project includes the extractor, FastAPI service, C# and Java client examples, a sanitized snapshot and focused checks.

Description previews are deliberately minimal. Each record includes the normalized description length and fingerprint. Staff administration and workflow metadata are excluded.

Read the project guide
python catalog_sample.py --live --limit 12
python catalog_sample.py --output replay
python -m uvicorn api:app --port 8765

GET /api/courses?q=knitting&limit=3
GET /api/quality
GET /openapi.json
Open API contractDownload quality evidence