# Catalog Trace

A focused Python + Playwright extraction sample using the real **RISD 2026–2027 Coursedog catalog**. It demonstrates public response discovery, URL discovery, HTML normalization, structured exports, validation, caching and a read-only REST interface.

This is an independent technical demonstration, not a past client commission. The saved capture contains **12 courses from the first results page**, not the entire catalog. The website does not crawl RISD when a visitor opens it.

## Run

Requires Python 3.12 or newer. From the extracted project directory:

```sh
python -m venv .venv
# Windows: .venv\Scripts\activate
# macOS/Linux: source .venv/bin/activate
python -m pip install -r requirements.txt
python -m playwright install chromium
python catalog_sample.py --live --limit 12
python -m uvicorn api:app --host 127.0.0.1 --port 8765
```

Open `http://127.0.0.1:8765`. An installed Edge browser can replace the downloaded Chromium: `python catalog_sample.py --live --browser-channel msedge --limit 12`. Each browser run uses a fresh isolated context.

For an offline, reproducible run with no source requests:

```sh
python catalog_sample.py --output replay
python -m unittest discover -s tests -v
```

The public fixture is a sanitized copy of the verified export. After a live run, a private local cache allows replay of the selected source records. Neither browser login nor a model API is required.

## Evidence and checks

`site/data/quality-report.json` records capture time, observed request method, HTTP status, response fingerprint, robots policy, actual discovered links and export fingerprint. The extractor observes the catalog's normal public **POST search response**; it does not call account or administration endpoints.

Required fields, stable identifiers, allowed source URLs and credit ranges are validated. Conflicting duplicates stop publication; identical duplicates are removed. Missing credit values remain null and zero stays zero. Spreadsheet formula prefixes are escaped in CSV. Text is rendered without HTML injection.

Full descriptions are normalized locally but not redistributed. Public exports include a tiny excerpt, normalized character count and SHA-256 fingerprint. Staff, account and workflow fields are excluded. Source ownership remains with the university.

## Outputs

- `courses.csv`, `courses.json`, `courses.jsonl`, `courses.jsonld`
- `GET /api/courses?q=knitting&limit=3&offset=0`
- `GET /api/courses?format=jsonld`
- `GET /api/health`, `GET /api/quality`, `GET /openapi.json`
- C# and Java REST client examples in `clients/`

Local API uses FastAPI. The public Cloudflare Pages deployment serves the same saved dataset through a small JavaScript worker. Refreshing the data requires running the Python extractor and redeploying its exports; the public site remains available independently of the development laptop.

## Honest scope

One Coursedog catalog is demonstrated. Acalog support, broad multi-catalog discovery and integration into an existing client codebase are future work. C# is tested against the local API; the Java 11 example is included but has not been runtime-tested on this machine. No previous commercial education-data experience is claimed.

Source catalog: https://risd.coursedog.com/courses

Code: MIT license. Bundled font: SIL Open Font License, with its license included. University data is not covered by the code license.
