- Track trainset/train/val/test.jsonl (focused excerpts only, no multi-MB text)
so the dataset can be used directly without the ~5h rebuild.
- DATASET.md: record schema, both serializations, load snippets (plain Python +
HF datasets), a text->triples fine-tuning sketch, eval notes, provenance.
- .gitignore: keep only the 100s-of-MB intermediates (samples_full) ignored.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>