Independent Python + Playwright sample
From a course catalog
to data you can trust.
Discover the public course response, normalize its records and inspect every result against its source.
| Course | Credits | Checks |
|---|
No matching course. Try a different title or code.
A small pipeline.
A complete chain of evidence.
Built for the practical work of catalog discovery, extraction, transformation and debugging.
- Discover
Playwright opens the public Coursedog catalog and observes its course response and actual course links.
- Normalize
Python cleans HTML, preserves credit ranges and creates stable course identifiers. Missing values stay explicit.
- Validate
Invalid required fields and conflicting duplicates stop publication. Each exported record retains its source URL.
- Deliver
CSV, JSON, JSONL, JSON-LD, a read-only REST API and an offline snapshot. No model API is required.
Run it. Review it. Reproduce it.
The downloadable project includes the extractor, FastAPI service, C# and Java client examples, a sanitized snapshot and focused checks.
Description previews are deliberately minimal. Each record includes the normalized description length and fingerprint. Staff administration and workflow metadata are excluded.
Read the project guidepython catalog_sample.py --live --limit 12
python catalog_sample.py --output replay
python -m uvicorn api:app --port 8765
GET /api/courses?q=knitting&limit=3
GET /api/quality
GET /openapi.jsonOpen API contractDownload quality evidence