MW verbal-root inventory
publicSLP1 root + explicit verb_type (750 genuineroot + 1363 root)
A single curated entry point for Sanskrit computational linguistics: our openly-licensed derived datasets (downloadable), plus the external stacks and APIs the project builds on — what each does, how to call it, and its license.
Derived from the Cologne Digital Sanskrit Dictionaries and the DCS corpus. Public assets are CC BY-SA 4.0, downloadable from the kosha data releases. Rendered from datasets.json — the single machine-readable source.
SLP1 root + explicit verb_type (750 genuineroot + 1363 root)
SLP1 headword → root
DCS lemma id ↔ CDSL headword (SLP1), 81.4% linked
SLP1 headword + per-dict provenance flags
MW entry id → Heritage anchor (25,140 covered, 97.6% anchor-resolved)
SLP1 lemma; whole-corpus counts + per-period vectors + core-vocab coverage
SLP1 lemma x layer(wn|mw|semdom) x sense_id; count_all + sense_rank + lemma_share; wave-1.5 dispersion de-biasing: n_texts + dispersion_dp (Gries DP) + largest_text_share + count_adj + sense_rank_adj; wave-2 Renou genre-stratified: count_bal_uniform + count_nonsastra + top_genre + top_genre_share (companion dcs_text_genre.tsv); provenance=attested|estimated + confidence
SLP1 headword → paradigm token (335 tokens)
per-saying records, public domain
akshara 7,347 · varṇa 48 · conjunct 999; corpus-wide + per DCS time-slot
declension class (14, lemma-final heuristic) × token gender (Masc|Fem|Neut|?) × 24 case.number cells; each cell carries token count, distinct lemma_ids, segmentable count, top-8 attested unsandhied forms, top-6 surface endings, up to 3 corpus examples (sentence + text + ref)
root (VERB lemma string, frequency floor >=2 total VERB tokens) x display category x number (sg|du|pl) x person (1|2|3) for finite cells, plus 8 non-finite buckets (PPP, Abs, Inf, FPP, PresPart, FutPart, PartOther, Other); each cell carries its top-5 attested forms with corpus counts
5 vargas × 5 DCS time-slots; column shares + counts + Δ(I→V)
lemma; total occurrences; co-occurrence count; ranked collocate list per lemma
one file per source-passage id (`<id>_<n>--<m>.csv`), 245 files; each row = a candidate parallel passage with a GOOD/PARTLY/… match verdict
root,count — one file per verb class (1.csv…10.csv), roots that have attested corpus forms
root;count — one file per verb class (1.csv…10.csv), includes prefixed verb forms
lemma; frequency; collocates — one file per historical period (1.csv…7.csv)
one file per source-passage id, same naming scheme as dcs-parallel-passages-full; includes a split 7z archive member
verb form; frequency count, corpus-wide
CompDic.csv (compound headword list, 37,333), cmps.csv (per-compound member breakdown, 401,478), names.csv (168,880 attested compounds w/ frequency vector, already noted in DATA_LAYERS_CENSUS.md), parts.csv (2,667), verbx.csv (3,400, verbal compounds); cmp400000.csv is present but empty (0 bytes)
final_member (MW uttarapada, orthography-folded: anusvāra ṃ/ṁ, avagraha, @/- markup) -> dictionary side (mw_class ∈ UTTARAPADA/KRT_STEM_MEMBER, mw_first_members = distinct-first-member TYPE count) vs corpus side (corpus_compounds = distinct attested word-forms, corpus_tokens = summed DCS frequency, corpus_first_members) + divergence cols (overlap_first, mw_only_first, corpus_only_first) + corpus_status ∈ {final 6249, form_variant 1289, nonfinal_only 1252, absent 10387}
stem-pair co-occurrence rows, ID range 1–222342
scenario id (G-01..G-18 dev, P-01..P-06 test) → gold + accepted CDSL dictionary codes over a 44-code answer space
(dict, L, k1) locus per correction record; line numbers are batch-time only, never a join key
union-headwords SLP1 key ↔ DCS lemma_id; canonical concordance-core schema (anchor_type/anchor_id/anchor_key_slp1/corpus_locus/match_method/confidence/evidence_count); ASSERTED tiers xref 12,836 · exact 61,373 · floor 311; the lossy relaxed tier (2,171) is QUARANTINED to dict_corpus_relaxed_candidates.tsv after the golden sample found 3/3 semantically wrong — per-tier table in data/concordance/BUILD_REPORT.md
concordance-core schema (anchor_type=parallel-verse/anchor_id=para:<textId>:<col1>|<col2>/target_locus verbatim/match_method=GOOD|PARTLY/confidence/evidence_count); source verdict GOOD (13,862 exact) or PARTLY (139,183 partial, word-diff attached) from the export's own meter-to-meter matcher, NOT the Q1 SLP1-tier vocabulary — see data/concordance/PARALLEL_BUILD_REPORT.md
maṇḍala + sūkta + verse + pada_letter (a/b/c/d, blank if unlettered) → pratīka text + full raw citation string
dict code x tag -> count + per-1,000-entry rate (96 distinct tags, 1,496,157 entries)
IAST + SLP1 lemma; per-slot counts, cosine distances (5 slot pairs), freq-shift baseline, graded rank + binary flag (Vedic→Epic)
event_id -> date, dict, headword_iast, old/new IAST, edit-op trace, corrector, latency, evidence level; _typed adds empirical error typology, _final is the canonical enriched view
surface form; constituent-stem split; part count; total corpus frequency; per-period/text frequency tail (semicolon-delimited)
one distinct SLP1 character n-gram per line, all lengths, sorted by length; derived from sanhw1.txt (union headword list)
dict_a x dict_b (unordered pair) -> shared / union / Jaccard, on the union's exact SLP1 <k1> key
neutral-model case id -> one MDF record; markers per data/schema/mdf-export-profile.json (C&G 2000 App. B order)
neutral-model case id -> one LIFT entry per file
Level A: (ak_varga_id, semdom_code/guid) many-to-many; Level B: AK eid -> semdom codes (top-6 candidates) / adjudicated gold code
SLP1 headword; 3x3 strata (E26 rank band x MW top-level sense count), seed 730; <=5 DCS attestation sentences per headword
one object per DCS chapter (slug); sentences[] each carry the sandhied text + ordered tokens (form, lemma, slp1, /w card href, upos, morph, gloss). 439 tokens, 434 (98.9%) linked to a kosha card via the concordance-core exact/floor tiers; 5 unlinked = DCS causative/denominative -ay stems + 1 indeclinable (honest residue, see reading/BUILD_REPORT.md). DCS is CC BY 4.0; this derivative ships BY-SA with the public tier.
one row per unique content-word lemma (SLP1), {rank, slp1, deva, iast, gloss}, core_rank-ordered
one row per analysed word (9,092; all 18 adhyayas); 21 fields = verse·lemma·devanagari·iast·form_type·code·tense·pada·vclass·root·root_tr·prefix·stem_end·gender·compound(TP/BV/DV)·mark·rule·sandhi·verse_iast·gloss_en·gloss_ru. Hand-curated (Combined sheet of SanskritGrammar/Concordance/Gita.xlsm); the garbled private-use Russian TRANSLITERATION column is dropped, the clean Cyrillic gloss kept (MG 13-07-2026).
one row per analysed word (9,091); structured morphology decoded from the Gita.xlsm Grammar-sheet shorthand (col AB) via the workbook's Abbreviations legend: verse·widx·form·lemma·root·pos·case·number·gender·person·tense·voice·nonfinite·derivation·compound·raw_morph. Nominals get case·number·gender; finite verbs person·number·tense·voice; participles/derivatives tagged; compound type TP/BV/DV/KD. raw_morph preserves the source shorthand.
DIVERGE (360) + GAP (919) rows from the W4 QA: gold Gita nominal case·number·gender NOT reproduced by kosha's hybrid inflections. Cols verse·form·lemma·gold_case_num_gender·class·kosha_analyses. Finding: DIVERGE 71% pronouns (confirms E1's flagged pronominal mis-modeling with attested text), GAP mostly compounds.
one viewer pack per adhyaya (gita-1..18.js, window.READING_DATA[slug]); 701 verses / 9,092 words, each word links to its kosha /w/ card (~99.5% linked). Devanagari + IAST + English + Russian gloss (EN/Русский toggle) per verse. Built from the W0 master gita_gold_master.tsv by scripts/build_reading_pack_gita.py. Each word also carries its sandhi rule (hover).
one row per distinct sandhi rule (161) attested across the whole Gita (3,412 junctions): rule (e.g. 'aḥ a → o ''), category (visarga/anusvara/vowel-coalescence/consonant), count, pct, example words+verse. Corpus-attested + frequency-ranked, aggregated from the master's per-word sandhi field. Teaching page reading/sandhi/index.html; per-word sandhi also shown on hover in the Gita reader.
one row per sandhi rule (2,181), ordered as a graded syllabus: rank, rule, category, count, pct, cumulative_pct, family_size, priority, lesson (10 lessons). Priority = frequency x class x environment-generality (MG ruling 14-07-2026); weights in data/sandhi/difficulty_weights.json (tunable). Learn 23 rules -> read 50% of all corpus sandhi; 79 -> 80%; 132 -> 90%. Teaching page reading/sandhi/curriculum/index.html.
one row per core-vocabulary lemma (6,667 of 7,120 Leonchenko core_rank lemmas; 453 dropped for having no committed dictionary card), ordered by core_rank: rank, lemma_slp1, deva, iast, gloss, core_rank, coverage_pct, cumulative_pct, lesson (134 lessons of 50), card_href. cumulative_pct is computed here (running sum of the source coverage_pct marginal weight), not copied from the source feed. Learn 284 lemmas -> read 30% of the core-vocabulary corpus mass; 1122 -> 50%; 4978 -> 70%. Teaching page reading/vocabulary/curriculum/index.html.
one item per row: aspect=vocabulary, type (recognition: word->gloss / recall: gloss->word), prompt, answer, distractors (same-lesson-band lemmas), rank (core_rank), evidence (docs/cards/<token>.json), source_dataset=vocab-curriculum (ARCHITECTURE shared item schema). 2 items per curriculum lemma (6,667 x 2 = 13,334). Also shipped as data/frequency/vocab_curriculum.apkg (Anki deck, genanki, one Basic note per lemma: front=deva/iast, back=gloss).
one row per (varga, lemma): varga_id, kanda, theme_iast, theme_slp1, semdom_keywords (A58 crosswalk cross-reference, up to 3 closest SIL semdom domain names, orientation only), lemma_slp1, deva, iast, gloss, core_rank, card_href. Sorted by varga file-order then core_rank (most frequent first) within each theme. 2,961 of 8,353 Amarakosa (varga, lemma) pairs across the 20 genuinely thematic vargas (the 4 grammatical/misc annexes excluded) kept -- the rest had no committed H947 vocab_curriculum.tsv row (no dead /w/ links).
one item per row: aspect=vocabulary-thematic, type (recognition: word->gloss / recall: gloss->word), prompt, answer, distractors (SAME-theme lemmas, not same-frequency-band), theme, varga_id, rank (core_rank), evidence (vocab_curriculum.tsv card_href), source_dataset=thematic-vocabulary (ARCHITECTURE shared item schema). 2 items per kept lemma (2,961 x 2 = 5,922). Also shipped as data/frequency/thematic_vocabulary.apkg (Anki deck, genanki, one sub-deck per theme).
one row per (lemma, declension/conjugation model) paradigm that has at least one corpus-attested cell (7,134 of 5,985 core-vocabulary lemmas x their models), ordered by class bucket (a-stems -> other-vowel-stems -> consonant-stems -> pronouns -> present-class-verbs -> other per drill_weights.json) then by lemma core_rank: rank, lemma_slp1, lemma_iast, model, kind (nominal/verb), bucket, core_rank, n_attested_cells, corpus_count, pct, cumulative_pct, lesson (10 lessons). Learn 4,862 paradigms -> cover 50% of attested nominal/verbal tokens; 6,351 -> 80%; 6,708 -> 90% (bucket-first ordering trades coverage-efficiency for pedagogical class sequencing — an honest, expected tradeoff, not a defect). Teaching page reading/morphology/curriculum/index.html.
one item per row: aspect=morphology, type (fill: lemma+cell->form / match: form->cell), prompt, answer, choices (4-way MCQ, same-lemma-paradigm distractors), distractors, rank (core_rank), evidence (dcs:<sent_id> locus), source_dataset=morphology-drills, corpus_count (ARCHITECTURE shared item schema). 12,000 items over the top 6,000/38,782 attested cells by (core_rank, corpus frequency) — --max-drill-cells 0 rebuilds the full set; the curriculum TSV is never scoped by this cap. Also shipped as data/morphology/morphology_drills.apkg (Anki deck, genanki, 6,000 fill-type notes). Web quiz reading/morphology/drills/index.html (self-contained MCQ, theme-aware, fill/match type filter).
one item per row: id, type (join/split/identify), rule, category, lesson, difficulty, question, answer, choices (4-way MCQ with same-class distractors), context (attested example sentence). 396 items over the 132 curriculum rules covering 90% of corpus sandhi (lessons 1-9; lesson-10 long tail excluded by default, --max-lesson to widen). Also shipped as data/sandhi/sandhi_drills.tsv (flat fallback) and data/sandhi/sandhi_drills.apkg (Anki deck, genanki). Web quiz reading/sandhi/drills/index.html (self-contained MCQ, theme-aware).
one row per distinct sandhi rule (13,012) aggregated across 41 DCS texts (707,936 sandhi events): rule, category, global_count, global_pct, n_texts, top_texts, examples. Induced by method A from DCS gold word-splits (96.3% Gita-gold frequency-mass coverage). The top 82 rules cover 80% of all corpus sandhi occurrences — the graded-curriculum backbone.
one row per unique DCS lemma (629, deduped from 717 Whitney-hub roots with corpus attestation -- 74 homonym-shared lemmas collapsed to avoid triple-counting the same corpus mass): rank, root_iast (one or more Whitney root strings joined by ' / ' when homonym-shared), dcs_lemma, grammar_class, dcs_status, attested_count, coverage_pct (cumulative), top_attested_forms.
101 hand-written etymological/explanatory notes on selected Gita words (verse·widx·form·lemma·root·etymology), from the Gita.xlsm Grammar sheet col AG (which the Combined master drops), aligned by verse+word-index. Curated highlight layer, not a full etymological dictionary.
148 verb roots attested in the Gita + 69 preverb-modified senses: root·preverb·combined·sense·count (empty preverb = base sense; √vac speak -> pra-vac declare; √gam go -> ava-gam understand). From the Gita.xlsm verbs sheet; a compositional dimension the Cologne dictionaries lack. Browsable page reading/upasarga/.
208 curated correct analyses (form_slp1·lemma_slp1·gcase·number·gender·source) for Sanskrit pronouns (sarvanaman), from the gold Gita attested pronoun forms. Applied to inflections as source='curated-gita-pronoun' via build_db.py --stage pronoun (non-destructive INSERT OR IGNORE). Fixed the W4 QA's pronominal mis-modeling: nominal agreement 93.0%->98.7%, divergences 360->73, gaps 919->588.
one row per gold-verified compound (759; from the 815-row gita-morphology-gold compound column, minus 56 ambiguous dual-tag rows), ordered as a graded syllabus: rank, lesson (1-4 = KD/TP/BV/DV, transparent types first per the MG 14-07-2026 ruling), type, type_name, compound, form, verse, corpus_freq, cumulative_freq_pct. Within a lesson, ranked by corpus frequency (VisualDCS Kompozity names.csv join). Companion data/samasa/reference.tsv (per-type ranked lookup) and data/samasa/samasa_drills.json (3,565 identify/split practice items: 759 gold-verified identify, 806 gold-verified split, 2,000 corpus-derived unverified-type split, capped from a 168,421-item ranked corpus pool per data/samasa/drill_weights.json). Teaching pages reading/samasa/{curriculum,drills,reference}/index.html.
one row per scored reading pack, ascending difficulty: order, slug, difficulty (composite in [0,1]), vocab, sandhi, morphology, compound (four per-axis loads in [0,1]), content_tokens, tokens, sentences, title. Composite = weighted sum of the four axes; weights live in data/difficulty/difficulty_weights.json (tunable -- a human should confirm, VERIFICATION R5). Companion reading_pack_difficulty.json (same rows + the weights) and the graded-reading page reading/difficulty/index.html (easiest first). Method + limitations: data/difficulty/METHODS.md.
one row per Gita chapter pack, ascending difficulty: order, slug, difficulty (reduced composite in [0,1]), vocab, sandhi, compound (three per-axis loads), content_tokens, tokens, sentences, title. The Gita packs carry no UD morphology, so they are scored on THREE axes -- vocab (slp1 -> lemma_frequency, non-compound content lemmas), sandhi (fraction of tokens carrying the pack's own per-token induced junction rule -- a real sandhi signal, NOT the 4-axis boundary proxy), compound (hyphen-lemma share) -- with the morphology weight dropped and the rest renormalised. Companion gita_reading_pack_difficulty.json + a labelled section on reading/difficulty/index.html.
one row per reading-pack sentence: pack, locus, dcs, syllables, metre, metre_type (vrtta/jati), method, confidence, text-preview. Two-tier + null method: strict vrtta via vidyut.chandas (method=vidyut-chandas, confidence=high, requires >=8 syllables); anustubh via a syllable heuristic (method=syllable-heuristic, confidence=medium -- every such tag is a whole-pada multiple of 8 in [8,32], and 840/840 land at exactly 16 = a half-sloka); everything else left unresolved with an empty metre (prose, headings, fragments -- never guessed). Companion metre_coverage.tsv (per-pack distribution + identified %).
vendored verbatim from vidyut-data/chandas/meters.tsv: meter name, type (vrtta), G/L pattern string. Consumed by vidyut.chandas.Chandas to classify strict vrttas. Vendored so kosha's metre annotator is self-contained (no sibling-clone dependency at build time).
one row per distinct '<upos>|<morph>' form signature over all ~5.69M DCS tokens: signature, count, share_pct. The signature is keyed exactly as the reading packs display it (build_reading_pack.morph_str), so a pack token joins with no re-derivation. Drives the difficulty scorer's morphology axis (rarer form = higher parsing load).
one object per DCS chapter (slug); sentences[] each carry the sandhied text + ordered tokens (form, lemma, slp1, /w card href, upos, morph, gloss). 392 tokens, 385 (98.2%) linked to a kosha card via concordance-core exact/floor tiers.
one object per DCS chapter (slug); sentences[] carry sandhied text + ordered tokens (form, lemma, slp1, /w href, upos, morph, gloss). 351 tokens, 337 (96.0%) linked.
one object per DCS chapter (slug); sentences[] carry sandhied text + ordered tokens (form, lemma, slp1, /w href, upos, morph, gloss). 900 tokens, 885 (98.3%) linked.
one object per DCS chapter (slug); sentences[] carry sandhied text + ordered tokens (form, lemma, slp1, /w href, upos, morph, gloss). 876 tokens, 854 (97.5%) linked.
H### id → verdict (ARCHIVE_AS_DONE/RERUN_NEEDED/RETIRE_AS_DEAD/RUNNABLE_NOW/GENUINELY_BLOCKED/NEEDS_STUB_FILL/UNCLEAR) + classifier verdict + adversarial-verifier verdict (upheld/refuted/uncontested)
kosha.db generated inflected forms (forms, include_heritage=False; 426,410 non-heritage rows) joined to DCS attested surface forms (dcs_full.sqlite token.form; 381,413 distinct) on form_key() equality — the length-preserving floor tier, promoted to exact where SLP1 keys are byte-identical (no NFD+strip path, D6/1b-2). Three buckets: AG 401,368 (of 426,410 generated) · G¬A 25,042 · A¬G 2 (of 381,413 attested). Canonical concordance-core schema (anchor_type/anchor_id/anchor_key_slp1/target_locus/match_method/confidence/evidence_count) + gen_source, tense_caveat, attested_form. Confidence from TIER_CONFIDENCE (exact 0.95 · floor 0.85). Loci are host-independent dcs:<sent_id>.
For every W1b AG-bucket row (lemma_slp1 + attested_form), resolves candidate grammatical cells from kosha.db inflections (nominal: gender/gcase/number where person IS NULL; verbal: model/tense/voice/person/number where person IS NOT NULL, v_p passive borrowing gaṇa from the same root's v_<gana> model per H855), derives every candidate cell with vidyut.prakriya (local library, no network — R12), and classifies the form: ok (exactly one distinct ordered-sūtra chain form_key-matches, exact 0.95 outranks floor 0.85 — TIER_CONFIDENCE, never a literal/1.0) · ambiguous (>1 distinct chain matches) · engine-error (candidate cells existed, none derivable) · no-derivation (no cell exists, or cells ran but none matched). derivation_chains.tsv is the (chain_id, step_index)->(source, sutra_code, step_result) sidecar the chain_id in derivation_status.tsv points into.
Inverts W2a's derivation_status.tsv (ok-status forms only — see notes) into one row per (sūtra, form, locus) triple: for every ok-status AG-bucket form, walks its ordered Ashtadhyayi-only sūtra chain (chain_id -> derivation_chains.tsv) and emits a concordance row per chain step. Canonical concordance-core schema (anchor_type=panini-sutra, anchor_id=sutra:<a.p.n>, anchor_key_slp1 empty, target_locus=dcs:<sent_id>[_<sub>], source_dataset=dcs, match_method in {exact,floor}, confidence from TIER_CONFIDENCE, evidence_count=1) + form_key_slp1, dcs_text, chain_position, chain_length, chain_id, derivation_status, tense_caveat.
one row per (reading pack, sentence, token index); joins each pack token to the three SanskritRussian PUBLIC site-tier layers — surface (token slp1 → surface_glossary), lemma (lemma slp1 → lemma_glossary), root (lemma → dcs_lemma2root → root_glossary) — emitting surface_ru/lemma_ru/root_ru + a layer_hit provenance column. 95.6% of tokens carry a lemma-layer RU gloss.
one object per curated saying, ordered by difficulty ascending: num/saying_id (Indische Sprüche numbering), deva, iast, translation_de, source_attribution, difficulty (W2a-reduced 2-axis composite in [0,1]), metre + metre_method + syllables (vidyut-chandas strict vṛtta else anuṣṭubh syllable heuristic), lines[].chunks[] with per-chunk sandhi-split tokens (t), per-seam X Y → Z junction rules (j, corpus-attested, data/sandhi/corpus_sandhi.tsv), per-token SLP1 lemma (lemma_slp1, H1312 — vidyut-cheda, honest null when unresolved) and per-token gloss_ru {surface,lemma,root} triple (H1312, null where uncovered), cross-boundary rules in lines[].xj. Grading spine: data/subhashita/subhashita_difficulty.tsv (all 7,537 sayings, full score decomposition). Curation: beginner_band.tsv + CURATION_NOTES.md (criteria + full reject log, no unlogged picks).
MANIFEST packs[] with pin_path + sha256; reading pins + sandhi L1-3 subsets + optional lemmas_for_srs.tsv
one row per Scherzl government relation (root/stem × case), verdict + evidence
PD work (normalised siglum-family -> IAST title) -> DCS 2021/2026 text; match_type in {complete, partial, absent}, DCS-anchored on the bounded 276-text inventory
one row per (headword slp1, numbered PWG sense_id, attestation); columns slp1/hom/sense_id/lemma/locus/cite/conf/method/rights/source/gloss/sent. Method tiers: ls 85,472 (PWG's own <ls> under the sense, conf 0.99, guaranteed-correct witness; MBh loci carry their resolved Nīlakaṇṭha-vulgate address) · locus 5 (DCS attestation verse-EQUAL to a sense's <ls>, conf 0.90 — canonically-numbered Vedic texts) · locus-mbh 48 (wave-1.5: DCS Mahābhārata attestation whose parvan+adhyāya matches a sense's <ls>-resolved vulgate adhyāya at ±1, conf 0.65–0.80, via the csl-atlas f8 crosswalk) · overlap 1,655 (shared proper-noun/Latin-binomial/digit gloss tokens across the DE/EN gap, conf 0.50–0.70). confidence<0.60 + unassigned residue parked in sense_review_queue.tsv, never dropped. A2 acceptance: <ls>-locus-resolution rate 99.3% on the 500-headword pilot; MBh <ls>→vulgate resolution 96% (7,055/7,353).
One row per sūtra in the named vidyut-0.4.0 Ashtadhyayi enumeration (Data.load_sutras / sutrapatha.tsv, n=3983 exact). Columns: sutra_id, sutra_text_slp1, exemplar_forms, exemplar_loci, texts, mean_chain_position, status in {lit, dark-unattested, dark-out-of-scope, dark-engine-gap}, scope_justification. Ratio 221:55:3707:0 — three dark classes never collapsed (ARCHITECTURE §5, VERIFICATION 3a-3/3a-8).
one item per row: id, type (classify: headword->declension-class token / odd-one-out: 4 headwords, pick the one whose class differs), question, answer, choices (4-way MCQ), index_token, gender, stem_class, member_count, tags (ARCHITECTURE shared item schema). 3,434 items over 332 declension-class tokens (indeclinables excluded), member-count-descending sampled so high-yield paradigms dominate (PER_TOKEN_CAP=15 classify / 8 odd-one-out groups per token). Also shipped as data/zaliznyak/zaliznyak_drills.apkg (Anki deck, genanki) and data/zaliznyak/zaliznyak_paradigm_classes.tsv (slim committed class index: token/gender/stem_class/member_count/representative). Web quiz reading/zaliznyak/drills/index.html (self-contained MCQ, theme-aware, type/gender filter).
one row per (pilot lemma, PWG sense_id) with MW/Apte inventory columns on the first PWG sense row only (not sense-aligned). Columns: lemma_slp1, hom, pwg_sense_id/gloss/loci, mw_sense_id/gloss, apte_sense_id/gloss, confidence, note. Scope = H1455 pilot 500 only.
one row per prose block (verse_label may be single 1.12 or range 1.4-6); verse_keys expands ranges; text = joined interlinear form+(gloss) lines from Gita.xlsm Prose sheet
one row per tracked PWG/PWK literary source, keyed on ls_code (the abbreviation, presentation marks stripped). 43 columns: status + status_gloss, citation_count + citation_count_safe + citation_count_provenance + citation_count_full, total_pages, volunteer, started/finished/index_posted/public_link dates, coordinating issue repo+number+URL, scan_dir + canonical spelling + Pages URL + repo size, and the cross-validation set (in_pwgbib, dict_citations, count_ratio, ru_subset_refs, scan_wired). book_no is NOT unique - do not key on it.
one row per bibliography abbreviation, keyed on abbrev (bib_code carried alongside), in two dated tables: pwg_ls_counts_2024-09-11.tsv (2837 buckets, ALL = 739056) and pwg_ls_counts_current.tsv (2846 buckets, ALL = 799500). Columns: total, abbrev, bib_code, gloss. Two synthetic buckets, NUMBER and UNKNOWN, complete the partition. Each table has a sibling .meta.json with input sha256s, source commits, and the ALL/NUMBER/UNKNOWN totals.
heritage_ref_subset.tsv (slp1 -> DICO anchor + gloss SHA-256 + word count; NO gloss text) · judge_fr_<arm>.jsonl (slp1 -> cross-lingual adequacy 0-5, 5 arms x 333) · heritage_ref_per_item.tsv (arm x slp1 -> chrF vs MW/FR/multi-ref + token-F1) · heritage_ref_scores.json (summary, gates, paired MW-FR premium with bootstrap CI)
pwg.tm.v1 record_id; entry_id / sense_id / fragment_id; SHA-256 of source/target strings
IRI under https://w3id.org/sanskrit-lexicon/repwg/ ; entry = (key1, homonym); sense = entry + edition layer + sense_tag slug
parvan/adhyaya/shloka (Nīlakaṇṭha vulgate) + fitted per-parvan continuous index -> four-state verdict in {present, absent, unchecked} x {present, absent, unchecked}, plus BORI locus, critical_score and matched-half counts
subject:grammar-lab:<slug> with Type-D edges to whitney-sec / whitney-root / zalizniak-1975|1978|2004
dict (mw|ap90) + SLP1 lemma + Cologne unit id (MW <L> record / AP90 {@N@} segment)
(key1, sense_index) + sense_no; every row basis "per Böhtlingk–Roth's citations"; dated_works are ls_source_map.json sigla
one row per PWG entry keyed on the Cologne L number: L, vol (1..7), col — PWG's own <pc> 'vol-Spalte' key, exactly the page key Cologne's servepdf.php honours for a multi-volume dictionary (H839: '{vol}-{col:04d}'; a bare column silently serves volume 1)
Rights-encumbered or unbackuped local-only assets — listed for discovery; available on request as rights clear.
SLP1 surface key (~190k keys), per verse-pair alignment
Cyrillic surface form → SLP1 lemma(s); frequency floor ≥ 20; 4,597 matched headwords before the floor, 205,012 occurrences over 1,521 files; per-course profile across 42 courses
SLP1; surface 190,838 · lemma 40,370 · root 2,021; 87% token coverage
10 tables: 444,773 entries · 323,425 lemmas · 692,403 senses · 1,378,401 forms · 6,917,018 inflections · 185,803 heritage_anchor rows (24,549 anchor-resolved) · 760 stem_bridge (+ meta, sources)
5,688,416 tokens · 754,726 sentences · 180,176 lemmas · 270 texts
580,552 corpus lines · 152 sources, verse-aligned (FTS5 index included)
Heritage entry anchors; VH↔SLP1 bridge validated
heritage_dico_gloss (mw_key1 → Heritage anchor + FR gloss), heritage_forms_oracle_disagreements (form → Heritage/kosha lemma disagreement class), heritage_only_forms (form → lemma, Heritage-only coverage)
vulgate verse address 5.<sarga>.<verse> (southern vulgate, 68 sargas)
parallels 40,573,260 (method='stopword', run='2022-partial', structured columns only, no verse text) + parallel_text 102 + text_names 128 + _id_fullname 106
base.db 320,902 corpus lines / 146 sources (corpus minus MW+Apte) + dict.db 254,037 lines / 2 sources (MW+Apte); FTS5 + pack_meta/pack_sources
(key1, subcard) card -> one MDF record; senses -> \sn (observed PWG numbering) + \de RU (+ \ge EN when promoted)
bucketed per data_root.py layout: tm/ (canonical store, subcard-keyed jsonl, 11,603 rows @ 02-08-2026) · layers/ (NWS packed tar.gz, 168k files) · manifests/ · raws/ (per-key window payloads) · telemetry/ · gatelogs/ · parked/ · corpus/ (corpus-gate dictionaries)
one row per headword slp1: {stratum(tm|control), originals:{likh,mw,apte,mac,pwg} html fragments, mt:{mw_ru,apte_ru,pwg_ru} html fragments, provenance sha256 per fetched part}; length-preserving SLP1 keys
Call these, don't clone them. Rendered from external_tools.json.
Ambuda (ambuda-org)
Rust Sanskrit toolkit: the kosha FST lexicon, a prakriyā (derivation) generator, the cheda segmenter, and a chandas meter identifier. The paradigm/lemma engine behind our RU translation kits and the Zaliznyak grammar index.
Our relation: We call vidyut for declension/conjugation paradigms and PPP validation rather than hand-rolling paradigm generation; it produces the paradigm tokens in zaliznyak-grammar-index.
Ambuda (ambuda-org)
Open-source Sanskrit reading platform and library (proofread texts, integrated dictionary + parser lookups). The reading front-end sibling of vidyut, built by the same community.
Our relation: Reference open-source reader; shares the vidyut engine we already consume. Not a data dependency of kosha.
Gérard Huet, INRIA
The Heritage dictionary + morphology generator + segmenter: DICO hypertext dictionary, MW-aligned entry pages, frequency TSVs, and OCaml morphology banks. A morphology oracle for form → lemma resolution.
Our relation: We align MW ↔ Heritage entries (see the mw-heritage-crosswalk dataset). Pull from the GitHub mirror — sanskrit.inria.fr is Anubis bot-walled. Mirror data itself is LGPLLR-pending in our restricted tier; the crosswalk is ours.
Amba Kulkarni, University of Hyderabad
The Sanskrit Computational Linguistics stack: morphological analyzer, sandhi splitter/segmenter, compound processor, Pāṇinian dependency parser, Amarakośa semantic net, and Dhātupāṭha resources. Already consumes the Cologne dictionaries.
Our relation: We call the live JSON endpoints and cross-validate against them; we do not clone the GPL source. It reuses our dictionary data downstream.
Sebastian Nehrdich et al., UC Berkeley
AI-driven Sanskrit stack: machine translation, deep research with references, parallel-passage exploration, segmentation, and OCR. A GPU morphology supplier and translation service.
Our relation: We reuse their MT error taxonomy (wired into the pwg_ru/mw_ru QA judges) and are a prospective Kosh-API consumer/supplier. GPU-morphology supplier to csl-atlas.
Oliver Hellwig
A sandhi-split, morphologically and lexically analysed diachronic corpus of Sanskrit (~5.7M tokens, 270 texts) with per-token lemma, POS and morphology in CoNLL-U. The reference corpus for lemma frequency and attestation.
Our relation: Source of the dcs-cdsl-xref crosswalk and the kosha-lemma-frequency sidecar. We ingest the canonical CoNLL-U once — never re-parse.
University of Cologne / UZH (C-SALT)
Accented Rig-Veda with per-word morphology (Casaretto word-split, Lubotsky padapāṭha), aligned translations, and a query API. The validation set for Vedic accent.
Our relation: The accented-Vedic validation set that unblocks the Vedic-accent axis of the Zaliznyak index — bulk-export once, cross-validate.
University of Cologne (Cologne / sanskrit-lexicon)
The canonical digitised Sanskrit dictionaries (MW, PWG, AP90, and 40+ more) as XML/text with scan-anchored entries and a lookup API. The lexical bedrock the whole project is built on.
Our relation: kosha collapses every CDSL dictionary's entry for a headword onto one scan-anchored page; union-headwords, mw-roots and mw-etymology all derive from CDSL source.