कोश
Gasuns Sanskrit Dictionary · the Sanskrit data-hub

Sanskrit NLP Data & Tools Directory

A single curated entry point for Sanskrit computational linguistics: our openly-licensed derived datasets (downloadable), plus the external stacks and APIs the project builds on — what each does, how to call it, and its license.

92public datasets
14restricted (on request)
8external tools & stacks

Our datasets

Derived from the Cologne Digital Sanskrit Dictionaries and the DCS corpus. Public assets are CC BY-SA 4.0, downloadable from the kosha data releases. Rendered from datasets.json — the single machine-readable source.

MW verbal-root inventory

public

SLP1 root + explicit verb_type (750 genuineroot + 1363 root)

2,113 rows · 75.3 KB · TSV

WhitneyRootscsl-atlasRussianTranslation verb worklist

Headword→root derivation table (Pāṇinian, 10 dicts)

public

SLP1 headword → root

9,377 rows · 1.5 MB · TSV

WhitneyRootscsl-atlas

DCS lemma ↔ CDSL headword crosswalk

public

DCS lemma id ↔ CDSL headword (SLP1), 81.4% linked

15,902 rows · 618.5 KB · TSV

simple-searchcsl-atlasVisualDCS

Cross-dictionary union headword index

public

SLP1 headword + per-dict provenance flags

323,425 rows · 11.8 MB · TSV

kosha crosswalkSanskritSpellCheck (union=N tags)PWG→EN pilot (94,753 subset)

MW ↔ Heritage (INRIA) entry-level crosswalk

public

MW entry id → Heritage anchor (25,140 covered, 97.6% anchor-resolved)

185,803 rows · 3.1 MB · TSV

kosha ingest (H345, kosha.db heritage_anchor table + /api/v1/lemma `heritage` witness)csl-atlas witness column (queued)

DCS lemma frequency sidecar (whole-corpus + per-period)

public

SLP1 lemma; whole-corpus counts + per-period vectors + core-vocab coverage

83,277 rows · 4.7 MB · TSV

kosha evidence layerNagari teaching order

DCS per-sense frequency sidecar (3-layer: WN synset / MW sense / semdom; attested WordSem + estimated MFS)

public

SLP1 lemma x layer(wn|mw|semdom) x sense_id; count_all + sense_rank + lemma_share; wave-1.5 dispersion de-biasing: n_texts + dispersion_dp (Gries DP) + largest_text_share + count_adj + sense_rank_adj; wave-2 Renou genre-stratified: count_bal_uniform + count_nonsastra + top_genre + top_genre_share (companion dcs_text_genre.tsv); provenance=attested|estimated + confidence

116,788 rows · 15.0 MB · TSV

kosha cards (N in this sense / M for the lemma + estimated chip)word_page SSR/prerender sense-freq blockRussianTranslation mfs_baseline.py (corpus MFS)semdom-amarakosha-crosswalk A58

Zaliznyak-style compact grammar-token index over all PWG

public

SLP1 headword → paradigm token (335 tokens)

98,639 rows · 5.8 MB · TSV

pwg_ru nominal layerkosha P4 grammar tokensgramdict mapping

Böhtlingk Indische Sprüche subhāṣitas

public

per-saying records, public domain

7,537 rows · 6.9 MB · JSONL

PWG <ls> citation links

DCS akshara/varṇa/ligature frequency tables

public

akshara 7,347 · varṇa 48 · conjunct 999; corpus-wide + per DCS time-slot

8,394 rows · 910.2 KB · CSV

Nagari script-teaching order

DCS nominal paradigm — case × number grid per declension class

public

declension class (14, lemma-final heuristic) × token gender (Masc|Fem|Neut|?) × 24 case.number cells; each cell carries token count, distinct lemma_ids, segmentable count, top-8 attested unsandhied forms, top-6 surface endings, up to 3 corpus examples (sentence + text + ref)

14 rows · 620.0 KB · JSON

VisualDCS sanskrit_nominal_dashboard.html

DCS verbal paradigm — per-root attested finite + non-finite cells

public

root (VERB lemma string, frequency floor >=2 total VERB tokens) x display category x number (sg|du|pl) x person (1|2|3) for finite cells, plus 8 non-finite buckets (PPP, Abs, Inf, FPP, PresPart, FutPart, PartOther, Other); each cell carries its top-5 attested forms with corpus counts

7,689 rows · 3.5 MB · JSON

VisualDCS sanskrit_paradigm_trainer.html (browse grid + frequency-weighted flashcard deck, with JSON deck export)

Consonant series (varga) by DCS period + Cramér's V

public

5 vargas × 5 DCS time-slots; column shares + counts + Δ(I→V)

5 rows · 7.6 KB · CSV

GasunsDhatu 2026 §2.6 (Table 5 replacement, L7 fix)

DCS syntagmatic table — all corpus lemmas (Приложение 7)

public

lemma; total occurrences; co-occurrence count; ranked collocate list per lemma

82,799 rows · 25.7 MB · CSV

DCS parallel-passage alignments, full-size export (PARA/Polnorazmernye)

public

one file per source-passage id (`<id>_<n>--<m>.csv`), 245 files; each row = a candidate parallel passage with a GOOD/PARTLY/… match verdict

506,787 rows · 56.2 MB · CSV

DCS verb roots grouped by class (with attested forms)

public

root,count — one file per verb class (1.csv…10.csv), roots that have attested corpus forms

463 rows · 4.4 KB · CSV

DCS verb-form frequency by class, with prefixed forms

public

root;count — one file per verb class (1.csv…10.csv), includes prefixed verb forms

8,454 rows · 110.0 KB · CSV

DCS syntagmatic tables of frequent lexical cores, by historical period (Приложение 6)

public

lemma; frequency; collocates — one file per historical period (1.csv…7.csv)

19,076 rows · 22.9 MB · CSV

DCS parallel-passage alignments, 'stop-word' variant export (PARA/Stopovye)

public

one file per source-passage id, same naming scheme as dcs-parallel-passages-full; includes a split 7z archive member

1.3 GB · CSV

DCS verb-form frequency in corpus (preliminary)

public

verb form; frequency count, corpus-wide

106 rows · 8.0 KB · CSV

DCS compound (samāsa) dictionary derivation

public

CompDic.csv (compound headword list, 37,333), cmps.csv (per-compound member breakdown, 401,478), names.csv (168,880 attested compounds w/ frequency vector, already noted in DATA_LAYERS_CENSUS.md), parts.csv (2,667), verbx.csv (3,400, verbal compounds); cmp400000.csv is present but empty (0 bytes)

613,758 rows · 161.4 MB · CSV

MW uttarapada index × DCS Kompozity — dictionary productivity vs corpus attestation

public

final_member (MW uttarapada, orthography-folded: anusvāra ṃ/ṁ, avagraha, @/- markup) -> dictionary side (mw_class ∈ UTTARAPADA/KRT_STEM_MEMBER, mw_first_members = distinct-first-member TYPE count) vs corpus side (corpus_compounds = distinct attested word-forms, corpus_tokens = summed DCS frequency, corpus_first_members) + divergence cols (overlap_first, mw_only_first, corpus_only_first) + corpus_status ∈ {final 6249, form_variant 1289, nonfinal_only 1252, absent 10387}

19,177 rows · 845.0 KB · TSV

kosha samāsa trainer member_side/member_recall drill ranking (scripts/build_samasa_trainer.py, --mw-rank corpus, default since H1398 20-07-2026)

Sanskrit-stem co-occurrence table (full)

public

stem-pair co-occurrence rows, ID range 1–222342

353,351 rows · 34.6 MB · CSV

Which-dictionary routing shared-task benchmark (v1.0)

public

scenario id (G-01..G-18 dev, P-01..P-06 test) → gold + accepted CDSL dictionary codes over a 44-code answer space

24 rows · 5.5 KB · JSON

csl-guides shared-tasks page (/about/shared-tasks) + leaderboard

Correction loci — locus-resolved correction records across all csl-corrections change files (v1.0)

public

(dict, L, k1) locus per correction record; line numbers are batch-time only, never a join key

39,540 rows · 9.2 MB · TSV

csl-atlas correction feed (loci heatmap + radar axes, memo backlog #2)docs/img/ corrections-native viz via scripts/build_correction_viz.py

CDSL headword ↔ DCS corpus concordance (B1, concordance core)

public

union-headwords SLP1 key ↔ DCS lemma_id; canonical concordance-core schema (anchor_type/anchor_id/anchor_key_slp1/corpus_locus/match_method/confidence/evidence_count); ASSERTED tiers xref 12,836 · exact 61,373 · floor 311; the lossy relaxed tier (2,171) is QUARANTINED to dict_corpus_relaxed_candidates.tsv after the golden sample found 3/3 semantically wrong — per-tier table in data/concordance/BUILD_REPORT.md

74,520 rows · 6.4 MB · TSV

concordance/dict/ static viewerQ2-Q4 concordance program

Bloomfield-style parallel-passage concordance across the DCS corpus (B3)

public

concordance-core schema (anchor_type=parallel-verse/anchor_id=para:<textId>:<col1>|<col2>/target_locus verbatim/match_method=GOOD|PARTLY/confidence/evidence_count); source verdict GOOD (13,862 exact) or PARTLY (139,183 partial, word-diff attached) from the export's own meter-to-meter matcher, NOT the Q1 SLP1-tier vocabulary — see data/concordance/PARALLEL_BUILD_REPORT.md

153,045 rows · 30.7 MB · TSV

concordance/parallels/ static viewerQ3-Q4 concordance program

Bloomfield 1906 Vedic Concordance — direct Ṛgveda citations (pratīka index)

public

maṇḍala + sūkta + verse + pada_letter (a/b/c/d, blank if unlettered) → pratīka text + full raw citation string

36,680 rows · 2.7 MB · TSV

parallel_passage_verses.tsv bloomfield_pratika columnconcordance/parallels/ static viewer

Markup-tag frequency census (44 Cologne v02 dicts)

public

dict code x tag -> count + per-1,000-entry rate (96 distinct tags, 1,496,157 entries)

671 rows · 18.1 KB · TSV

SanskritLexicography FEATURES_INDEX (E39)Uprava DATA_LAYERS_CENSUS section 3markup QA / under-marking triage

Sanskrit lexical semantic change scores (PPMI-cosine, DCS 5-slot chronology)

public

IAST + SLP1 lemma; per-slot counts, cosine distances (5 slot pairs), freq-shift baseline, graded rank + binary flag (Vedic→Epic)

3,049 rows · 283.6 KB · TSV

paper A57 (LChange'27/NLP4DH candidate, ARTICLES.md)

Correction-event log, full trio (all/typed/final)

public

event_id -> date, dict, headword_iast, old/new IAST, edit-op trace, corrector, latency, evidence level; _typed adds empirical error typology, _final is the canonical enriched view

52,498 rows · 169.8 MB · CSV

SanskritLexicography FEATURES_INDEX (E41; release-only view already E32)Uprava DATA_LAYERS_CENSUS section 4corrector-behaviour / error-typology studies (A36 family)

DCS proper-name compound splits with frequency vectors (Kompozity names.csv)

public

surface form; constituent-stem split; part count; total corpus frequency; per-period/text frequency tail (semicolon-delimited)

168,879 rows · 86.5 MB · CSV

SanskritLexicography FEATURES_INDEX (E42)Uprava DATA_LAYERS_CENSUS section 4compound (samasa) morphology / onomastics studies

Pan-CDSL headword character-n-gram oracle (allngramtxt.txt)

public

one distinct SLP1 character n-gram per line, all lengths, sorted by length; derived from sanhw1.txt (union headword list)

6,656,616 rows · 78.5 MB · TXT

SanskritLexicography FEATURES_INDEX (F43)Uprava DATA_LAYERS_CENSUS section 4n-gram spell-check method (suspect-headword detection)

Headword pairwise-overlap matrix (15-dict union)

public

dict_a x dict_b (unordered pair) -> shared / union / Jaccard, on the union's exact SLP1 <k1> key

105 rows · 2.8 KB · TSV

SanskritLexicography FEATURES_INDEX (E40)Uprava DATA_LAYERS_CENSUS section 3A40 headword-inventory paper (H675)

CDSL MDF pilot export (50-case neutral-model slice, MW-focused)

public

neutral-model case id -> one MDF record; markers per data/schema/mdf-export-profile.json (C&G 2000 App. B order)

250 rows · 88.1 KB · MDF (SFM, ONE RECORD PER FILE)

Lexique Pro smoke test (H722)SIL MDF program (papers/SIL_MDF_ECOSYSTEM_CORRELATION.md)

CDSL LIFT pilot export (fourth serialization, H721)

public

neutral-model case id -> one LIFT entry per file

250 rows · 302.7 KB · LIFT (XML)

Webonary / Dictionary App Builder channel assessment (H726)

semdom <-> Amarakosha crosswalk (SIL semantic domains <-> AK vargas + synsets, ID pairs)

public

Level A: (ak_varga_id, semdom_code/guid) many-to-many; Level B: AK eid -> semdom codes (top-6 candidates) / adjudicated gold code

5,898 rows · 408.2 KB · CSV (LEVEL A VARGA MAP) + TSV (LEVEL B CANDIDATES, GOLD, ANNOTATOR PASSES)

csl-standards MDF/LIFT exports (sd semantic-domain field, H721)paper A58 (GWC / LT4HALA / eLex)

MW definition-generation eval — frozen 500-headword sample + attestations + scored baseline outputs (H730)

public

SLP1 headword; 3x3 strata (E26 rank band x MW top-level sense count), seed 730; <=5 DCS attestation sentences per headword

500 rows · 1.3 MB · TSV+JSONL

H730 protocol doc docs/DEFGEN_MW_GLOSS_EVAL_PROTOCOL.md (eLex/EURALEX/IJL paper candidate)

DCS reading pack — Nala 1 (Nalopākhyāna, Mahābhārata 3.50)

public

one object per DCS chapter (slug); sentences[] each carry the sandhied text + ordered tokens (form, lemma, slp1, /w card href, upos, morph, gloss). 439 tokens, 434 (98.9%) linked to a kosha card via the concordance-core exact/floor tiers; 5 unlinked = DCS causative/denominative -ay stems + 1 indeclinable (honest residue, see reading/BUILD_REPORT.md). DCS is CC BY 4.0; this derivative ships BY-SA with the public tier.

65 rows · 134.2 KB · JSON

reading/ static viewer

SRS deck — Rung B1 demo (Nala 1 content vocabulary, frequency-ordered)

public

one row per unique content-word lemma (SLP1), {rank, slp1, deva, iast, gloss}, core_rank-ordered

164 rows · 24.9 KB · JSON

Systema-Sanscriticum Saraswati SRS (Wave-2 importer, flag-gated)

Bhagavadgita gold word-by-word analysis (all 18 adhyayas)

public

one row per analysed word (9,092; all 18 adhyayas); 21 fields = verse·lemma·devanagari·iast·form_type·code·tense·pada·vclass·root·root_tr·prefix·stem_end·gender·compound(TP/BV/DV)·mark·rule·sandhi·verse_iast·gloss_en·gloss_ru. Hand-curated (Combined sheet of SanskritGrammar/Concordance/Gita.xlsm); the garbled private-use Russian TRANSLITERATION column is dropped, the clean Cyrillic gloss kept (MG 13-07-2026).

9,092 rows · 1.1 MB · TSV

reading/ Gita packs (W1)gita sandhi/morphology/root-preverb datasets (W2/W3/W6)E1 inflection-engine QA (W4)

Bhagavadgita gold morphology + compound dataset

public

one row per analysed word (9,091); structured morphology decoded from the Gita.xlsm Grammar-sheet shorthand (col AB) via the workbook's Abbreviations legend: verse·widx·form·lemma·root·pos·case·number·gender·person·tense·voice·nonfinite·derivation·compound·raw_morph. Nominals get case·number·gender; finite verbs person·number·tense·voice; participles/derivatives tagged; compound type TP/BV/DV/KD. raw_morph preserves the source shorthand.

9,091 rows · 587.1 KB · TSV

E1 inflection-engine QA (W4 / H874)

Gita inflection-engine QA — divergence ledger (E1 attested-corpus check)

public

DIVERGE (360) + GAP (919) rows from the W4 QA: gold Gita nominal case·number·gender NOT reproduced by kosha's hybrid inflections. Cols verse·form·lemma·gold_case_num_gender·class·kosha_analyses. Finding: DIVERGE 71% pronouns (confirms E1's flagged pronominal mis-modeling with attested text), GAP mostly compounds.

1,279 rows · 65.3 KB · TSV

E1 hybrid forms layer — candidate disputed/gap-fill corrections (human @DO)

Bhagavadgita reading packs — all 18 adhyayas (gold)

public

one viewer pack per adhyaya (gita-1..18.js, window.READING_DATA[slug]); 701 verses / 9,092 words, each word links to its kosha /w/ card (~99.5% linked). Devanagari + IAST + English + Russian gloss (EN/Русский toggle) per verse. Built from the W0 master gita_gold_master.tsv by scripts/build_reading_pack_gita.py. Each word also carries its sandhi rule (hover).

701 rows · 187.9 KB · JS

reading/ static viewer

Bhagavadgita sandhi — frequency-ranked rules

public

one row per distinct sandhi rule (161) attested across the whole Gita (3,412 junctions): rule (e.g. 'aḥ a → o ''), category (visarga/anusvara/vowel-coalescence/consonant), count, pct, example words+verse. Corpus-attested + frequency-ranked, aggregated from the master's per-word sandhi field. Teaching page reading/sandhi/index.html; per-word sandhi also shown on hover in the Gita reader.

161 rows · 14.2 KB · TSV

reading/ viewer (per-word sandhi hover)reading/sandhi/ page

Graded sandhi curriculum — ordered teaching syllabus

public

one row per sandhi rule (2,181), ordered as a graded syllabus: rank, rule, category, count, pct, cumulative_pct, family_size, priority, lesson (10 lessons). Priority = frequency x class x environment-generality (MG ruling 14-07-2026); weights in data/sandhi/difficulty_weights.json (tunable). Learn 23 rules -> read 50% of all corpus sandhi; 79 -> 80%; 132 -> 90%. Teaching page reading/sandhi/curriculum/index.html.

2,181 rows · 134.7 KB · TSV

reading/sandhi/curriculum/ pagegraded sandhi teaching (Phase 4)

Frequency-graded vocabulary curriculum — ordered teaching syllabus

public

one row per core-vocabulary lemma (6,667 of 7,120 Leonchenko core_rank lemmas; 453 dropped for having no committed dictionary card), ordered by core_rank: rank, lemma_slp1, deva, iast, gloss, core_rank, coverage_pct, cumulative_pct, lesson (134 lessons of 50), card_href. cumulative_pct is computed here (running sum of the source coverage_pct marginal weight), not copied from the source feed. Learn 284 lemmas -> read 30% of the core-vocabulary corpus mass; 1122 -> 50%; 4978 -> 70%. Teaching page reading/vocabulary/curriculum/index.html.

6,667 rows · 1.1 MB · TSV

reading/vocabulary/curriculum/ pagevocab-drills dataset

Vocabulary drills — recognition/recall practice items

public

one item per row: aspect=vocabulary, type (recognition: word->gloss / recall: gloss->word), prompt, answer, distractors (same-lesson-band lemmas), rank (core_rank), evidence (docs/cards/<token>.json), source_dataset=vocab-curriculum (ARCHITECTURE shared item schema). 2 items per curriculum lemma (6,667 x 2 = 13,334). Also shipped as data/frequency/vocab_curriculum.apkg (Anki deck, genanki, one Basic note per lemma: front=deva/iast, back=gloss).

13,334 rows · 7.2 MB · JSON

vocab_curriculum.apkg (Anki deck)Systema/csl-guides quiz component (future, via the shared item schema)

Thematic vocabulary axis — corpus vocabulary grouped by Amarakosa varga

public

one row per (varga, lemma): varga_id, kanda, theme_iast, theme_slp1, semdom_keywords (A58 crosswalk cross-reference, up to 3 closest SIL semdom domain names, orientation only), lemma_slp1, deva, iast, gloss, core_rank, card_href. Sorted by varga file-order then core_rank (most frequent first) within each theme. 2,961 of 8,353 Amarakosa (varga, lemma) pairs across the 20 genuinely thematic vargas (the 4 grammatical/misc annexes excluded) kept -- the rest had no committed H947 vocab_curriculum.tsv row (no dead /w/ links).

2,961 rows · 590.2 KB · TSV

reading/vocabulary/thematic/ pagethematic-vocab-drills dataset

Thematic vocabulary drills — recognition/recall items grouped by Amarakosa theme

public

one item per row: aspect=vocabulary-thematic, type (recognition: word->gloss / recall: gloss->word), prompt, answer, distractors (SAME-theme lemmas, not same-frequency-band), theme, varga_id, rank (core_rank), evidence (vocab_curriculum.tsv card_href), source_dataset=thematic-vocabulary (ARCHITECTURE shared item schema). 2 items per kept lemma (2,961 x 2 = 5,922). Also shipped as data/frequency/thematic_vocabulary.apkg (Anki deck, genanki, one sub-deck per theme).

5,922 rows · 3.5 MB · JSON

thematic_vocabulary.apkg (Anki deck)Systema/csl-guides quiz component (future, via the shared item schema)

Graded morphology curriculum — corpus-attested paradigms, ordered teaching syllabus

public

one row per (lemma, declension/conjugation model) paradigm that has at least one corpus-attested cell (7,134 of 5,985 core-vocabulary lemmas x their models), ordered by class bucket (a-stems -> other-vowel-stems -> consonant-stems -> pronouns -> present-class-verbs -> other per drill_weights.json) then by lemma core_rank: rank, lemma_slp1, lemma_iast, model, kind (nominal/verb), bucket, core_rank, n_attested_cells, corpus_count, pct, cumulative_pct, lesson (10 lessons). Learn 4,862 paradigms -> cover 50% of attested nominal/verbal tokens; 6,351 -> 80%; 6,708 -> 90% (bucket-first ordering trades coverage-efficiency for pedagogical class sequencing — an honest, expected tradeoff, not a defect). Teaching page reading/morphology/curriculum/index.html.

7,134 rows · 478.3 KB · TSV

reading/morphology/curriculum/ pagemorphology-drills dataset

Morphology drills — fill/match practice items, evidence-linked to a DCS corpus locus

public

one item per row: aspect=morphology, type (fill: lemma+cell->form / match: form->cell), prompt, answer, choices (4-way MCQ, same-lemma-paradigm distractors), distractors, rank (core_rank), evidence (dcs:<sent_id> locus), source_dataset=morphology-drills, corpus_count (ARCHITECTURE shared item schema). 12,000 items over the top 6,000/38,782 attested cells by (core_rank, corpus frequency) — --max-drill-cells 0 rebuilds the full set; the curriculum TSV is never scoped by this cap. Also shipped as data/morphology/morphology_drills.apkg (Anki deck, genanki, 6,000 fill-type notes). Web quiz reading/morphology/drills/index.html (self-contained MCQ, theme-aware, fill/match type filter).

12,000 rows · 7.3 MB · JSON

reading/morphology/drills/ page (web quiz)morphology_drills.apkg (Anki deck)Systema/csl-guides quiz component (future, via the shared item schema)

Sandhi drills — join/split/identify practice items

public

one item per row: id, type (join/split/identify), rule, category, lesson, difficulty, question, answer, choices (4-way MCQ with same-class distractors), context (attested example sentence). 396 items over the 132 curriculum rules covering 90% of corpus sandhi (lessons 1-9; lesson-10 long tail excluded by default, --max-lesson to widen). Also shipped as data/sandhi/sandhi_drills.tsv (flat fallback) and data/sandhi/sandhi_drills.apkg (Anki deck, genanki). Web quiz reading/sandhi/drills/index.html (self-contained MCQ, theme-aware).

396 rows · 207.4 KB · JSON

reading/sandhi/drills/ page (web quiz)sandhi_drills.apkg (Anki deck)graded sandhi teaching (Phase 4, surface 4b)

Corpus sandhi — frequency-ranked rules across 41 texts

public

one row per distinct sandhi rule (13,012) aggregated across 41 DCS texts (707,936 sandhi events): rule, category, global_count, global_pct, n_texts, top_texts, examples. Induced by method A from DCS gold word-splits (96.3% Gita-gold frequency-mass coverage). The top 82 rules cover 80% of all corpus sandhi occurrences — the graded-curriculum backbone.

13,012 rows · 3.2 MB · TSV

reading/sandhi/ pages (planned)graded sandhi curriculum (planned)

Roots frequency + attestation curriculum (W2b)

public

one row per unique DCS lemma (629, deduped from 717 Whitney-hub roots with corpus attestation -- 74 homonym-shared lemmas collapsed to avoid triple-counting the same corpus mass): rank, root_iast (one or more Whitney root strings joined by ' / ' when homonym-shared), dcs_lemma, grammar_class, dcs_status, attested_count, coverage_pct (cumulative), top_attested_forms.

629 rows · 116.5 KB · TSV

WhitneyRoots (a 'learn roots in this order' layer to import instead of re-deriving its own ranking)Systema (future roots drill ordering)

Bhagavadgita etymology notes

public

101 hand-written etymological/explanatory notes on selected Gita words (verse·widx·form·lemma·root·etymology), from the Gita.xlsm Grammar sheet col AG (which the Combined master drops), aligned by verse+word-index. Curated highlight layer, not a full etymological dictionary.

101 rows · 12.2 KB · TSV

Gita reader / lexical enrichment

Sanskrit root x preverb (upasarga) semantics

public

148 verb roots attested in the Gita + 69 preverb-modified senses: root·preverb·combined·sense·count (empty preverb = base sense; √vac speak -> pra-vac declare; √gam go -> ava-gam understand). From the Gita.xlsm verbs sheet; a compositional dimension the Cologne dictionaries lack. Browsable page reading/upasarga/.

214 rows · 6.2 KB · TSV

reading/upasarga/ page/w/ root-card upasarga panel (app/word_page.py)

Curated pronoun-paradigm corrections (E1/W4 fix)

public

208 curated correct analyses (form_slp1·lemma_slp1·gcase·number·gender·source) for Sanskrit pronouns (sarvanaman), from the gold Gita attested pronoun forms. Applied to inflections as source='curated-gita-pronoun' via build_db.py --stage pronoun (non-destructive INSERT OR IGNORE). Fixed the W4 QA's pronominal mis-modeling: nominal agreement 93.0%->98.7%, divergences 360->73, gaps 919->588.

208 rows · 8.5 KB · TSV

kosha inflections layer (source=curated-gita-pronoun)E1 hybrid forms

Samasa (compound) analysis trainer

public

one row per gold-verified compound (759; from the 815-row gita-morphology-gold compound column, minus 56 ambiguous dual-tag rows), ordered as a graded syllabus: rank, lesson (1-4 = KD/TP/BV/DV, transparent types first per the MG 14-07-2026 ruling), type, type_name, compound, form, verse, corpus_freq, cumulative_freq_pct. Within a lesson, ranked by corpus frequency (VisualDCS Kompozity names.csv join). Companion data/samasa/reference.tsv (per-type ranked lookup) and data/samasa/samasa_drills.json (3,565 identify/split practice items: 759 gold-verified identify, 806 gold-verified split, 2,000 corpus-derived unverified-type split, capped from a 168,421-item ranked corpus pool per data/samasa/drill_weights.json). Teaching pages reading/samasa/{curriculum,drills,reference}/index.html.

759 rows · 50.8 KB · TSV

reading/samasa/{curriculum,drills,reference}/ pagessamasa_drills.apkg (Anki deck)csl-guides samasa-quiz (cross-linked, not duplicated)graded compound-analysis teaching (pedagogy Wave 1, W1c)

Reading-pack difficulty scores + graded ordering (W2a)

public

one row per scored reading pack, ascending difficulty: order, slug, difficulty (composite in [0,1]), vocab, sandhi, morphology, compound (four per-axis loads in [0,1]), content_tokens, tokens, sentences, title. Composite = weighted sum of the four axes; weights live in data/difficulty/difficulty_weights.json (tunable -- a human should confirm, VERIFICATION R5). Companion reading_pack_difficulty.json (same rows + the weights) and the graded-reading page reading/difficulty/index.html (easiest first). Method + limitations: data/difficulty/METHODS.md.

5 rows · 607 B · TSV

reading/difficulty/ page (graded reading sequence)graded-reader auto-levelling (W2a join point, ARCHITECTURE data-flow)W3a metre-in-reading (H951, consumes the ordered packs)

Gita reading-pack difficulty -- reduced 3-axis ordering (W2a follow-up)

public

one row per Gita chapter pack, ascending difficulty: order, slug, difficulty (reduced composite in [0,1]), vocab, sandhi, compound (three per-axis loads), content_tokens, tokens, sentences, title. The Gita packs carry no UD morphology, so they are scored on THREE axes -- vocab (slp1 -> lemma_frequency, non-compound content lemmas), sandhi (fraction of tokens carrying the pack's own per-token induced junction rule -- a real sandhi signal, NOT the 4-axis boundary proxy), compound (hyphen-lemma share) -- with the morphology weight dropped and the rest renormalised. Companion gita_reading_pack_difficulty.json + a labelled section on reading/difficulty/index.html.

18 rows · 1.8 KB · TSV

reading/difficulty/ page (Gita section)graded-reader auto-levelling within the Gita (once UD morphology is added, these fold into the 4-axis score)

Reading-pack per-verse metre annotation (W3a)

public

one row per reading-pack sentence: pack, locus, dcs, syllables, metre, metre_type (vrtta/jati), method, confidence, text-preview. Two-tier + null method: strict vrtta via vidyut.chandas (method=vidyut-chandas, confidence=high, requires >=8 syllables); anustubh via a syllable heuristic (method=syllable-heuristic, confidence=medium -- every such tag is a whole-pada multiple of 8 in [8,32], and 840/840 land at exactly 16 = a half-sloka); everything else left unresolved with an empty metre (prose, headings, fragments -- never guessed). Companion metre_coverage.tsv (per-pack distribution + identified %).

1,095 rows · 143.3 KB · TSV

SanskritKaraoke (metre trainer UI -- the consuming surface; kosha emits the data layer only)reading packs (per-verse metre for graded reading, field 3.9)

vidyut-chandas metre definitions (vendored)

public

vendored verbatim from vidyut-data/chandas/meters.tsv: meter name, type (vrtta), G/L pattern string. Consumed by vidyut.chandas.Chandas to classify strict vrttas. Vendored so kosha's metre annotator is self-contained (no sibling-clone dependency at build time).

144 rows · 5.1 KB · TSV

scripts/build_reading_pack_metre.py (reading-pack-metre)

DCS morphological-form frequency table (difficulty-scorer signal)

public

one row per distinct '<upos>|<morph>' form signature over all ~5.69M DCS tokens: signature, count, share_pct. The signature is keyed exactly as the reading packs display it (build_reading_pack.morph_str), so a pack token joins with no re-derivation. Drives the difficulty scorer's morphology axis (rarer form = higher parsing load).

840 rows · 27.6 KB · TSV

scripts/build_difficulty_scorer.py (morphology axis)

DCS reading pack — Nala 2 (Nalopakhyana, Mahabharata 3.51)

public

one object per DCS chapter (slug); sentences[] each carry the sandhied text + ordered tokens (form, lemma, slp1, /w card href, upos, morph, gloss). 392 tokens, 385 (98.2%) linked to a kosha card via concordance-core exact/floor tiers.

61 rows · 119.8 KB · JSON

reading/ static viewerreading-pack-difficulty (W2a scorer)

DCS reading pack — Nala 3 (Nalopakhyana, Mahabharata 3.52)

public

one object per DCS chapter (slug); sentences[] carry sandhied text + ordered tokens (form, lemma, slp1, /w href, upos, morph, gloss). 351 tokens, 337 (96.0%) linked.

51 rows · 103.4 KB · JSON

reading/ static viewerreading-pack-difficulty (W2a scorer)

DCS reading pack — Hitopadesa, Prastavika (opening)

public

one object per DCS chapter (slug); sentences[] carry sandhied text + ordered tokens (form, lemma, slp1, /w href, upos, morph, gloss). 900 tokens, 885 (98.3%) linked.

125 rows · 269.0 KB · JSON

reading/ static viewerreading-pack-difficulty (W2a scorer)

DCS reading pack — Kiratarjuniya 1 (Bharavi)

public

one object per DCS chapter (slug); sentences[] carry sandhied text + ordered tokens (form, lemma, slp1, /w href, upos, morph, gloss). 876 tokens, 854 (97.5%) linked.

92 rows · 261.3 KB · JSON

reading/ static viewerreading-pack-difficulty (W2a scorer)

Handoff lifecycle gold standard (170 adversarially-verified verdicts)

public

H### id → verdict (ARCHIVE_AS_DONE/RERUN_NEEDED/RETIRE_AS_DEAD/RUNNABLE_NOW/GENUINELY_BLOCKED/NEEDS_STUB_FILL/UNCLEAR) + classifier verdict + adversarial-verifier verdict (upheld/refuted/uncontested)

170 rows · 311.3 KB · JSONL

handoff-lifecycle sweep tool agreement measurement (H1251)Uprava handoff-status audit baseline

Generated↔attested morphology audit (A3, W1b)

public

kosha.db generated inflected forms (forms, include_heritage=False; 426,410 non-heritage rows) joined to DCS attested surface forms (dcs_full.sqlite token.form; 381,413 distinct) on form_key() equality — the length-preserving floor tier, promoted to exact where SLP1 keys are byte-identical (no NFD+strip path, D6/1b-2). Three buckets: AG 401,368 (of 426,410 generated) · G¬A 25,042 · A¬G 2 (of 381,413 attested). Canonical concordance-core schema (anchor_type/anchor_id/anchor_key_slp1/target_locus/match_method/confidence/evidence_count) + gen_source, tense_caveat, attested_form. Confidence from TIER_CONFIDENCE (exact 0.95 · floor 0.85). Loci are host-independent dcs:<sent_id>.

401,368 rows · 32.1 MB · TSV

A4 derivation capture (W2a)

vidyut-prakriya derivation harness over the AG bucket (A4, W2a)

public

For every W1b AG-bucket row (lemma_slp1 + attested_form), resolves candidate grammatical cells from kosha.db inflections (nominal: gender/gcase/number where person IS NULL; verbal: model/tense/voice/person/number where person IS NOT NULL, v_p passive borrowing gaṇa from the same root's v_<gana> model per H855), derives every candidate cell with vidyut.prakriya (local library, no network — R12), and classifies the form: ok (exactly one distinct ordered-sūtra chain form_key-matches, exact 0.95 outranks floor 0.85 — TIER_CONFIDENCE, never a literal/1.0) · ambiguous (>1 distinct chain matches) · engine-error (candidate cells existed, none derivable) · no-derivation (no cell exists, or cells ran but none matched). derivation_chains.tsv is the (chain_id, step_index)->(source, sutra_code, step_result) sidecar the chain_id in derivation_status.tsv points into.

401,368 rows · 25.9 MB · TSV

A4 sūtra->corpus inversion (W2b DONE, H1390 — data/concordance/paninian_concordance.tsv)W3a sūtra-coverage map (H1468 DONE — data/concordance/sutra_coverage_map.tsv)data-v0.3.0 release asset

Pāṇinian sūtra ↔ DCS corpus concordance (A4, W2b)

public

Inverts W2a's derivation_status.tsv (ok-status forms only — see notes) into one row per (sūtra, form, locus) triple: for every ok-status AG-bucket form, walks its ordered Ashtadhyayi-only sūtra chain (chain_id -> derivation_chains.tsv) and emits a concordance row per chain step. Canonical concordance-core schema (anchor_type=panini-sutra, anchor_id=sutra:<a.p.n>, anchor_key_slp1 empty, target_locus=dcs:<sent_id>[_<sub>], source_dataset=dcs, match_method in {exact,floor}, confidence from TIER_CONFIDENCE, evidence_count=1) + form_key_slp1, dcs_text, chain_position, chain_length, chain_id, derivation_status, tense_caveat.

893,482 rows · 82.8 MB · TSV

concordance/panini/index.html web viewer (kwic_<adhyaya>.js shards)W3a sūtra-coverage map (H1468 DONE)data-v0.3.0 release asset

Inline Sanskrit→Russian gloss layer (reading packs, W-RU-a)

public

one row per (reading pack, sentence, token index); joins each pack token to the three SanskritRussian PUBLIC site-tier layers — surface (token slp1 → surface_glossary), lemma (lemma slp1 → lemma_glossary), root (lemma → dcs_lemma2root → root_glossary) — emitting surface_ru/lemma_ru/root_ru + a layer_hit provenance column. 95.6% of tokens carry a lemma-layer RU gloss.

2,958 rows · 190.0 KB · TSV

reading/data/*.json reading packs (additive token.gloss_ru)reading/index.html Russian gloss toggle

Beginner subhāṣita reader pack — graded Indische Sprüche band (W-RU-b + gloss.ru)

public

one object per curated saying, ordered by difficulty ascending: num/saying_id (Indische Sprüche numbering), deva, iast, translation_de, source_attribution, difficulty (W2a-reduced 2-axis composite in [0,1]), metre + metre_method + syllables (vidyut-chandas strict vṛtta else anuṣṭubh syllable heuristic), lines[].chunks[] with per-chunk sandhi-split tokens (t), per-seam X Y → Z junction rules (j, corpus-attested, data/sandhi/corpus_sandhi.tsv), per-token SLP1 lemma (lemma_slp1, H1312 — vidyut-cheda, honest null when unresolved) and per-token gloss_ru {surface,lemma,root} triple (H1312, null where uncovered), cross-boundary rules in lines[].xj. Grading spine: data/subhashita/subhashita_difficulty.tsv (all 7,537 sayings, full score decomposition). Curation: beginner_band.tsv + CURATION_NOTES.md (criteria + full reject log, no unlogged picks).

106 rows · 533.7 KB · JSON

reading/subhashita/ page (graded beginner reader, difficulty-ordered, RU-gloss toggle)data/subhashita/subhashita_beginner_anki.apkg (Anki export of the band, GlossRu back field)reading/RU_GLOSS_COVERAGE.md (this pack's row, folded in alongside the 5 DCS reading packs)

Start chteniya cohort pack freeze (Hitopadesa-0 + subhashita-beginner + sandhi L1-3)

public

MANIFEST packs[] with pin_path + sha256; reading pins + sandhi L1-3 subsets + optional lemmas_for_srs.tsv

5 rows · 950.6 KB · JSON+TSV

Systema-Sanscriticum resources/data/cohort_start_chteniya/ (H2106)reading_pack multi-pack import (H2110)

Scherzl case-government relations × DCS-treebank adjudication

public

one row per Scherzl government relation (root/stem × case), verdict + evidence

1,168 rows · 88.5 KB · TSV

Sanskrit parser / valence-frame stackBuhlerLeitfaden_1923 government_lexicon (corpus-confirmation layer)

Poona Dictionary × DCS corpus coverage crosswalk (H1336)

public

PD work (normalised siglum-family -> IAST title) -> DCS 2021/2026 text; match_type in {complete, partial, absent}, DCS-anchored on the bounded 276-text inventory

118 rows · 6.5 KB · TSV

csl-atlas /tools/pd-dcs-coverage Observable pagereports/PD_DCS_CORPUS_COVERAGE_2026.md

Per-sense corpus attestation (नागदन्त layer, H1455 wave-1)

public

one row per (headword slp1, numbered PWG sense_id, attestation); columns slp1/hom/sense_id/lemma/locus/cite/conf/method/rights/source/gloss/sent. Method tiers: ls 85,472 (PWG's own <ls> under the sense, conf 0.99, guaranteed-correct witness; MBh loci carry their resolved Nīlakaṇṭha-vulgate address) · locus 5 (DCS attestation verse-EQUAL to a sense's <ls>, conf 0.90 — canonically-numbered Vedic texts) · locus-mbh 48 (wave-1.5: DCS Mahābhārata attestation whose parvan+adhyāya matches a sense's <ls>-resolved vulgate adhyāya at ±1, conf 0.65–0.80, via the csl-atlas f8 crosswalk) · overlap 1,655 (shared proper-noun/Latin-binomial/digit gloss tokens across the DE/EN gap, conf 0.50–0.70). confidence<0.60 + unassigned residue parked in sense_review_queue.tsv, never dropped. A2 acceptance: <ls>-locus-resolution rate 99.3% on the 500-headword pilot; MBh <ls>→vulgate resolution 96% (7,055/7,353).

87,092 rows · 14.6 MB · TSV

concordance/senses/ static viewer (sense-sharded KWIC)kosha-sense-frequency (H1453) — shared attestation→sense assignment, two witnesses

Aṣṭādhyāyī sūtra-coverage / dark-class map (A4, W3a)

public

One row per sūtra in the named vidyut-0.4.0 Ashtadhyayi enumeration (Data.load_sutras / sutrapatha.tsv, n=3983 exact). Columns: sutra_id, sutra_text_slp1, exemplar_forms, exemplar_loci, texts, mean_chain_position, status in {lit, dark-unattested, dark-out-of-scope, dark-engine-gap}, scope_justification. Ratio 221:55:3707:0 — three dark classes never collapsed (ARCHITECTURE §5, VERIFICATION 3a-3/3a-8).

3,983 rows · 1.0 MB · TSV

A4 exit check / research reportsconcordance/panini coverage view (W4a)data-v0.3.0 release asset

Zaliznyak declension-class drills — classify a headword or spot the odd one out

public

one item per row: id, type (classify: headword->declension-class token / odd-one-out: 4 headwords, pick the one whose class differs), question, answer, choices (4-way MCQ), index_token, gender, stem_class, member_count, tags (ARCHITECTURE shared item schema). 3,434 items over 332 declension-class tokens (indeclinables excluded), member-count-descending sampled so high-yield paradigms dominate (PER_TOKEN_CAP=15 classify / 8 odd-one-out groups per token). Also shipped as data/zaliznyak/zaliznyak_drills.apkg (Anki deck, genanki) and data/zaliznyak/zaliznyak_paradigm_classes.tsv (slim committed class index: token/gender/stem_class/member_count/representative). Web quiz reading/zaliznyak/drills/index.html (self-contained MCQ, theme-aware, type/gender filter).

3,434 rows · 1.8 MB · JSON

reading/zaliznyak/drills/ page (web quiz)zaliznyak_drills.apkg (Anki deck)

Pilot cross-dictionary sense view (PWG · MW · Apte) — 500 headwords

public

one row per (pilot lemma, PWG sense_id) with MW/Apte inventory columns on the first PWG sense row only (not sense-aligned). Columns: lemma_slp1, hom, pwg_sense_id/gloss/loci, mw_sense_id/gloss, apte_sense_id/gloss, confidence, note. Scope = H1455 pilot 500 only.

7,359 rows · 2.6 MB · TSV

concordance/senses/crossdict.html (side-by-side pilot viewer)

Bhagavadgita interlinear prose paraphrase (Prose sheet)

public

one row per prose block (verse_label may be single 1.12 or range 1.4-6); verse_keys expands ranges; text = joined interlinear form+(gloss) lines from Gita.xlsm Prose sheet

653 rows · 222.8 KB · TSV

reading/index.html Prose togglereading/data/gita_prose.js

PWG literary-source scan-indexing campaign registry (2025-2026)

public

one row per tracked PWG/PWK literary source, keyed on ls_code (the abbreviation, presentation marks stripped). 43 columns: status + status_gloss, citation_count + citation_count_safe + citation_count_provenance + citation_count_full, total_pages, volunteer, started/finished/index_posted/public_link dates, coordinating issue repo+number+URL, scan_dir + canonical spelling + Pages URL + repo size, and the cross-validation set (in_pwgbib, dict_citations, count_ratio, ru_subset_refs, scan_wired). book_no is NOT unique - do not key on it.

82 rows · 37.3 KB · TSV

csl-observatory reports/pwg_scan_index.mdcsl-observatory /scan-index dashboard pagedata/pwg_scan_index_tracker/pwg_etext_candidate_queue.tsv (e-text extraction ranking, H1715)data/pwg_scan_index_tracker/scan_target_audit.tsv (resolver wiring, H1714)data/pwg_scan_index_tracker/pwg_citation_count_provenance.tsv (citation-count provenance + the denominator contract, H2874)

PWG <ls> citation counts per bibliography abbreviation (work-family rollup)

public

one row per bibliography abbreviation, keyed on abbrev (bib_code carried alongside), in two dated tables: pwg_ls_counts_2024-09-11.tsv (2837 buckets, ALL = 739056) and pwg_ls_counts_current.tsv (2846 buckets, ALL = 799500). Columns: total, abbrev, bib_code, gloss. Two synthetic buckets, NUMBER and UNKNOWN, complete the partition. Each table has a sibling .meta.json with input sha256s, source commits, and the ALL/NUMBER/UNKNOWN totals.

5,683 rows · 179.7 KB · TSV

csl-observatory data/pwg_scan_index_tracker/pwg_citation_count_provenance.tsvcsl-observatory data/pwg_scan_index_tracker/pwg_scan_index.tsv (citation_count_safe / _full)csl-observatory reports/pwg_citation_count_provenance.mdcsl-observatory /scan-index dashboard page (dictionary-wide coverage denominator)

Definition-generation eval: Heritage (Huet) French glosses as an independent second reference (H2408)

public

heritage_ref_subset.tsv (slp1 -> DICO anchor + gloss SHA-256 + word count; NO gloss text) · judge_fr_<arm>.jsonl (slp1 -> cross-lingual adequacy 0-5, 5 arms x 333) · heritage_ref_per_item.tsv (arm x slp1 -> chrF vs MW/FR/multi-ref + token-F1) · heritage_ref_scores.json (summary, gates, paired MW-FR premium with bootstrap CI)

3,742 rows · 265.3 KB · TSV + JSONL + JSON

docs/DEFGEN_HERITAGE_SECOND_REFERENCE_EVAL.md (report of record: reference-invariant arm ranking + MW-familiarity premium +0.13..+0.25)docs/DEFGEN_MW_GLOSS_EVAL_PROTOCOL.md next-step #4 (closed by this dataset)

PWG German–Russian translation memory (canonical v1 four-format pack)

public

pwg.tm.v1 record_id; entry_id / sense_id / fragment_id; SHA-256 of source/target strings

2,392 rows · 97.0 MB · JSONL + TMX + TEI-XML + TURTLE

RussianTranslation TM interchangecsl-standards LLOD modelling (generator stays in RussianTranslation)

PWG DE edition graph — OntoLex-Lemon + TEI Lex-0 sidecars

public

IRI under https://w3id.org/sanskrit-lexicon/repwg/ ; entry = (key1, homonym); sense = entry + edition layer + sense_tag slug

11,581 rows · 44.2 MB · TURTLE + TEI-XML + JSON

csl-standards LLOD modellingRussianTranslation SPARQL queries (release/query/*.rq)edition-graph consumers needing PWG/PW/SCH/PWKVN/NWS sense relations

MBh Nīlakaṇṭha vulgate × BORI critical-edition presence verdicts (H2845)

public

parvan/adhyaya/shloka (Nīlakaṇṭha vulgate) + fitted per-parvan continuous index -> four-state verdict in {present, absent, unchecked} x {present, absent, unchecked}, plus BORI locus, critical_score and matched-half counts

83,971 rows · 5.5 MB · CSV

SanskritLexicography RussianTranslation ls_links.MbhEtext — renders E / E† beside every MBh scan link on a card (PR #1753)csl-atlas data/forensic/MBH_ETEXT_PRESENCE_CENSUS.md

Grammar Lab Wave-1 topic graph (Whitney + Zalizniak root alternation / verbal morphology)

public

subject:grammar-lab:<slug> with Type-D edges to whitney-sec / whitney-root / zalizniak-1975|1978|2004

32 rows · 414.9 KB · JSON + TSV + YML

Systema-Sanscriticum Grammar Lab import (H2493 G2)

MW/AP90 sense units vs the PWG-family store: matched / unalignable / absent-candidate verdicts (wave 4)

public

dict (mw|ap90) + SLP1 lemma + Cologne unit id (MW <L> record / AP90 {@N@} segment)

2,767 rows · 820.6 KB · JSONL

SanskritLexicography edition-relations roadmap (issue #1736)csl-atlas A09 sense-alignment paper

PWG per-sense attestation window (Ceiling C2 phase 1)

public

(key1, sense_index) + sense_no; every row basis "per Böhtlingk–Roth's citations"; dated_works are ls_source_map.json sigla

53,003 rows · 17.3 MB · JSONL

pwg_ru cards (era badge)Ceiling C2 phase 2 curated dating tableC7 residue repair queue

PWG entry L → printed volume/column (print-scan anchor table, H3457)

public

one row per PWG entry keyed on the Cologne L number: L, vol (1..7), col — PWG's own <pc> 'vol-Spalte' key, exactly the page key Cologne's servepdf.php honours for a multi-volume dictionary (H839: '{vol}-{col:04d}'; a bare column silently serves volume 1)

122,730 rows · 1.4 MB · TSV

app/word_page_ux.py print-scan anchors (H3457 staging word pages) — rebuilds the PWG scan URL through kosha.scan_resolver with the volumeany card-level consumer of docs/cards/*.json scan_url for PWG — every committed PWG scan_url (48,540 at 25-08-2026) is bare-page and must be overlaid with this table until the cards are regenerated

Restricted & in preparation

Rights-encumbered or unbackuped local-only assets — listed for discovery; available on request as rights clear.

Sa→Ru word-aligned corpus lexicon

restricted

SLP1 surface key (~190k keys), per verse-pair alignment

1,093,391 rows · 277.1 MB · JSONL

SanskritRussian glossarypwg_ru TM
available on request / in preparation

Sanskrit terminology glossary of the устный корпус (taught-term frequency)

restricted

Cyrillic surface form → SLP1 lemma(s); frequency floor ≥ 20; 4,597 matched headwords before the floor, 205,012 occurrences over 1,521 files; per-course profile across 42 courses

1,513 rows · 115.4 KB · TSV

samskrte.ru ₽5,000 membership tier (candidate)
available on request / in preparation

Ranked Sa→Ru glossary (surface/lemma/root layers)

restricted

SLP1; surface 190,838 · lemma 40,370 · root 2,021; 87% token coverage

233,229 rows · 147.0 MB · JSONL

live search sitepwg_ru gates
available on request / in preparation

kosha unified lookup database

restricted

10 tables: 444,773 entries · 323,425 lemmas · 692,403 senses · 1,378,401 forms · 6,917,018 inflections · 185,803 heritage_anchor rows (24,549 anchor-resolved) · 760 stem_bridge (+ meta, sources)

444,773 rows · 1.6 GB · SQLITE

kosha APIP2 static cache
available on request / in preparation

DCS 2026 full corpus database

restricted

5,688,416 tokens · 754,726 sentences · 180,176 lemmas · 270 texts

5,688,416 rows · 878.2 MB · SQLITE

VisualDCS dashboardskosha frequencyakshara stats
available on request / in preparation

SamudraManthanam Sa↔Ru verse-parallel corpus

restricted

580,552 corpus lines · 152 sources, verse-aligned (FTS5 index included)

580,552 rows · 624.8 MB · SQLITE

pwg_ru gate dictsmw_ru validationcorpus-lexicon upstream
available on request / in preparation

Sanskrit Heritage (INRIA) local mirror

restricted

Heritage entry anchors; VH↔SLP1 bridge validated

MIXED (DICO HYPERTEXT, TSV, OCAML BANKS)

mw-heritage-crosswalkHeadwordLists coverage studies
available on request / in preparation

Heritage (INRIA) form-level crosswalk extras (gloss, disagreements, forms-only)

restricted

heritage_dico_gloss (mw_key1 → Heritage anchor + FR gloss), heritage_forms_oracle_disagreements (form → Heritage/kosha lemma disagreement class), heritage_only_forms (form → lemma, Heritage-only coverage)

1,037,239 rows · 457.9 MB · TSV

kosha ingest (H111, kosha.db forms table: heritage_only_forms.tsv -> 951,991 source='heritage' rows, 928,262 distinct forms)kosha default-off enforcement (H696, R7 ruling 10-07-2026: heritage witnesses opt-in only via ?heritage=1; excluded from the static tier)
available on request / in preparation

Sundarakāṇḍa two-tier commentary apparatus (Leonov LP volume)

restricted

vulgate verse address 5.<sarga>.<verse> (southern vulgate, 68 sargas)

1,955 rows · 2.9 MB · JSON

sundarakanda_print_master (data/book/, MD+DOCX)per-sarga interactive apparatus HTML (data/apparatus/)
available on request / in preparation

DCS parallel-passage stop-word run, queryable DB (M9 D2)

restricted

parallels 40,573,260 (method='stopword', run='2022-partial', structured columns only, no verse text) + parallel_text 102 + text_names 128 + _id_fullname 106

40,573,260 rows · 10.3 GB · SQLITE

available on request / in preparation

SamudraManthanam offline app packs (base + dict)

restricted

base.db 320,902 corpus lines / 146 sources (corpus minus MW+Apte) + dict.db 254,037 lines / 2 sources (MW+Apte); FTS5 + pack_meta/pack_sources

574,939 rows · 293.9 MB · SQLITE

SamudraManthanam mobile/offline app
available on request / in preparation

PWG->RU translation cards as MDF (bilingual \de RU / \ge EN records)

restricted

(key1, subcard) card -> one MDF record; senses -> \sn (observed PWG numbering) + \de RU (+ \ge EN when promoted)

2,350 rows · 3.7 MB · MDF (SFM, BLANK-LINE-DELIMITED RECORDS, SINGLE FILE)

Lexique Pro / Webonary / Dictionary App Builder (SIL toolchain)SIL MDF program (papers/SIL_MDF_ECOSYSTEM_CORRELATION.md)
available on request / in preparation

PWG→RU nonstop-pipeline working set (TM store, layers, manifests, raws, gate logs)

restricted

bucketed per data_root.py layout: tm/ (canonical store, subcard-keyed jsonl, 11,603 rows @ 02-08-2026) · layers/ (NWS packed tar.gz, 168k files) · manifests/ · raws/ (per-key window payloads) · telemetry/ · gatelogs/ · parked/ · corpus/ (corpus-gate dictionaries)

11,603 rows · 683.1 MB · GIT REPO + LFS (JSONL/TAR.GZ/JSON)

RussianTranslation nonstop lanes (PC / samskrte.ru / routine)bounded_staged_run.py --data-rootspot_check_daily.py / lane_guard.py / digest_daily.py
available on request / in preparation

akshara.ru MT benchmark pilot corpus (H3455 lane A)

restricted

one row per headword slp1: {stratum(tm|control), originals:{likh,mw,apte,mac,pwg} html fragments, mt:{mw_ru,apte_ru,pwg_ru} html fragments, provenance sha256 per fetched part}; length-preserving SLP1 keys

304 rows · 49.3 MB · JSONL (PARSED) + HTML (RAW)

H3456 blinded pwg_ru-vs-akshara-MT benchmarkfuture SanskritRussian mw_ru/apte_ru comparison
available on request / in preparation

External tools & stacks

Call these, don't clone them. Rendered from external_tools.json.

vidyut

external

Ambuda (ambuda-org)

Rust Sanskrit toolkit: the kosha FST lexicon, a prakriyā (derivation) generator, the cheda segmenter, and a chandas meter identifier. The paradigm/lemma engine behind our RU translation kits and the Zaliznyak grammar index.

Our relation: We call vidyut for declension/conjugation paradigms and PPP validation rather than hand-rolling paradigm generation; it produces the paradigm tokens in zaliznyak-grammar-index.

Ambuda

external

Ambuda (ambuda-org)

Open-source Sanskrit reading platform and library (proofread texts, integrated dictionary + parser lookups). The reading front-end sibling of vidyut, built by the same community.

Our relation: Reference open-source reader; shares the vidyut engine we already consume. Not a data dependency of kosha.

Call / docs ↗ AGPL-3.0 (platform); texts vary

Gérard Huet, INRIA

The Heritage dictionary + morphology generator + segmenter: DICO hypertext dictionary, MW-aligned entry pages, frequency TSVs, and OCaml morphology banks. A morphology oracle for form → lemma resolution.

Our relation: We align MW ↔ Heritage entries (see the mw-heritage-crosswalk dataset). Pull from the GitHub mirror — sanskrit.inria.fr is Anubis bot-walled. Mirror data itself is LGPLLR-pending in our restricted tier; the crosswalk is ours.

Amba Kulkarni, University of Hyderabad

The Sanskrit Computational Linguistics stack: morphological analyzer, sandhi splitter/segmenter, compound processor, Pāṇinian dependency parser, Amarakośa semantic net, and Dhātupāṭha resources. Already consumes the Cologne dictionaries.

Our relation: We call the live JSON endpoints and cross-validate against them; we do not clone the GPL source. It reuses our dictionary data downstream.

DharmaMitra

external

Sebastian Nehrdich et al., UC Berkeley

AI-driven Sanskrit stack: machine translation, deep research with references, parallel-passage exploration, segmentation, and OCR. A GPU morphology supplier and translation service.

Our relation: We reuse their MT error taxonomy (wired into the pwg_ru/mw_ru QA judges) and are a prospective Kosh-API consumer/supplier. GPU-morphology supplier to csl-atlas.

Call / docs ↗ Free web service; models vary — see site

Oliver Hellwig

A sandhi-split, morphologically and lexically analysed diachronic corpus of Sanskrit (~5.7M tokens, 270 texts) with per-token lemma, POS and morphology in CoNLL-U. The reference corpus for lemma frequency and attestation.

Our relation: Source of the dcs-cdsl-xref crosswalk and the kosha-lemma-frequency sidecar. We ingest the canonical CoNLL-U once — never re-parse.

Call / docs ↗ CC BY 4.0

VedaWeb

external

University of Cologne / UZH (C-SALT)

Accented Rig-Veda with per-word morphology (Casaretto word-split, Lubotsky padapāṭha), aligned translations, and a query API. The validation set for Vedic accent.

Our relation: The accented-Vedic validation set that unblocks the Vedic-accent axis of the Zaliznyak index — bulk-export once, cross-validate.

Call / docs ↗ CC BY 4.0

University of Cologne (Cologne / sanskrit-lexicon)

The canonical digitised Sanskrit dictionaries (MW, PWG, AP90, and 40+ more) as XML/text with scan-anchored entries and a lookup API. The lexical bedrock the whole project is built on.

Our relation: kosha collapses every CDSL dictionary's entry for a headword onto one scan-anchored page; union-headwords, mw-roots and mw-etymology all derive from CDSL source.

Call / docs ↗ CC BY-SA 4.0