SarvaVeda LLM ← Home
Plan of Action

Plan of Action

The plan divided into 160 numbered units across four stages; the sources the corpus is assembled from; and an honest account of what is loaded, what is queued, and what has still to be acquired.

॥ अनन्ता वै वेदाः ॥
Section 1

The plan in four stages

The work divides into four stages. They are thematic rather than strictly sequential — collection continues while the lexicon is built, and the lexicon deepens while the applications are written.

Stage 1 Collect and load the texts

Veda, Vedāṅga, Upāṅga and Upaveda; the Purāṇas, Itihāsa, Kāvya and the ancillary literature, together with their commentaries. Where a text is not already proof-read, it enters through the digitisation route.

Loading is not mere ingestion. The Bhāṣya is processed so that its Vyākaraṇa and śāstra references are linked to their targets; the related portions are brought together, so a Mantra or Brāhmaṇa passage stands beside its Pada-pāṭha and its Bhāṣyam; and the MAP analysis is produced for each mantra on demand rather than stored. Throughout, a master list is maintained that can absorb further texts and references without being restructured.

Stage 2 Pada Kośa — the vocabulary

Three vocabularies, built by different means because they are different in kind.

Rūḍha — the conventional vocabulary, drawn from the kośas. Being non-compositional by definition, it must be enumerated rather than derived.

Vaidika — every available Pada-pāṭha ingested, Pada-pāṭha generated where none exists, and the unique words with their meanings extracted from the Veda Bhāṣyam. A Pada-pāṭha processor brings uniformity at load time and keeps the processed padas separately, so that the Krama and Vikṛti sandhi layers can be built upon them.

Yaugika — the Vyākaraṇa Vocabulary Engine: the Pāṇinian word-space, over a hundred million forms, each carrying its complete sūtra-by-sūtra derivation, with meanings in Telugu, Kannada, Hindi and English. Telugu leads, because the team's native expertise gives the subtlest sense.

Alongside these: Nirukta etymology, with traditional nirvacana set beside the formal Pāṇinian vyutpatti; the liṅga rules from Nāma-liṅgānuśāsana, Mahābhāṣya and Kāśikā; Krama and the eight Vikṛti forms from Jaṭā to Ghana, with rule-verified sandhi and svara at every junctura; and Varṇa Krama — akṣara-by-akṣara decomposition with svara and mātrā, built directly on Śikṣā and Prātiśākhya.

Stage 3 Synthesis and applications

Data synthesised across the verticals to establish Vākya analysis, and the cross-references that run through Veda, Vedāṅga, Upāṅga and Upaveda connected so the corpus reads as one body rather than many.

Anuvāda — translation of the mantras into Indian languages founded solely on the Bhāṣya; where no Bhāṣya exists, on the principles of Sāyaṇācārya. Samānatā — an engine determining similarity in śabda and in artha.

On these rest the applications: Śrauta Prayoga generation with every step carrying a traceable sūtra citation; the Mīmāṃsā Nyāya engine; Sāmaveda stotra generation with viṣṭuti patterns and stobha insertion; Tarka sentence structure in the śābda-bodha style with pariṣkāra; Sāma sound analysis, audio in and notation out, with the Ūha engine that re-frames melody onto new chandas; and the Chatur-Darśana engine, which interprets one sentence in parallel through Tarka, Vyākaraṇa, Mīmāṃsā and Vedānta — the śāstrārtha assembly, computed.

Stage 4 Presentation

The work reaches its audience: integration with general search; the Vedic perspective offered on any topic on demand, with societal reference; APIs for researchers seeking Vedic references with plain meanings; video, audio and web publication; webinars; the VVS textbooks; short-form material across thousands of topics; a daily presence; and a considered śāstra perspective on contemporary events.

Section 2

Where the work stands

The plan is deliberately granular. Each unit carries a permanent number, a definition of done, and its dependencies — so progress is a matter of record rather than impression.

160numbered work units
15complete
26ready to start
227corpus artefacts held
2.9 GBof primary text

Unit numbers (SVU-001 onward) never change and never encode a stage, so work can be resequenced without renumbering. The same discipline governs every identifier in the project: an identifier records what a thing is, never where it currently sits.

Section 3

The working groups

The work is owned by group, not unit by unit, so new work inherits a group instead of waiting to be allocated. Each group has a convener who is answerable for its units. The Tech Team is not a further group — it carries the units whose deliverable is software.

Group Working group Convener Units Complete
G1Book Collection & Source ManagementVKG93
G2Data Processing, Analysis & IngestionChandra Chaganti195
G3Pada Kośa ReviewVKG333
G4Specialist Tasks — Audio ResearchNatraj Kavuri110
G5Mīmāṃsā / Vedavākya CategorizationVijay Krishna200
G6TestingGayathri Ramasubramanian30
G7Review — Book Review & User InteractionVishwanath30
G8Presentation / Multimedia (MMP)Chandra Chaganti70
TechTech TeamTech Team Lead554
On membership. Only conveners are named here. The full membership of each group is recorded on the Foundation's internal roster; it is not published on this page.
Section 4

Data sources

The corpus is assembled from the Foundation's own digitised holdings, the open Sanskrit lexical corpora, and a Pāṇinian derivation engine that generates grammatical forms rather than storing them.

Source What it contributes Items Size
Veda mūlaSaṃhitā text of the four Vedas, with svara3086 MB
Pada-pāṭhaWord-separated recitational text77 MB
BhāṣyaTraditional commentary — Sāyaṇa, Bhaṭṭa Bhāskara and others942,324 MB
VedāṅgaŚikṣā, Prātiśākhya, Kalpa and the ancillary śāstras63233 MB
PrayogaRitual manuals and performance texts33321 MB
Lexical — Cologne Digital Sanskrit Lexicon44 dictionaries: Monier-Williams, Apte, Amarakośa, Vācaspatyam, Śabdakalpadruma and others44
Lexical — indic-dict collections19 further kośa builds used as independent cross-check witnesses19
Grammatical enginePāṇinian derivation engine with the Dhātupāṭha — 2,229 dhātus, 5,160 sūtras, 128 kṛt and 181 taddhita pratyayas1
Pāṇinian reference corporaAṣṭādhyāyī commentary and annotation sets surveyed for reuse4
On the lexical sources. The lexical collections are held for comparison and corroboration: of the distinct Sanskrit words assembled, roughly two-thirds are attested by more than one independent kośa. Where a word rests on a single witness it is recorded as such rather than presented as settled.
Section 5

The corpus, division by division

Every division of the corpus, what is held in it, and what the seventy-two volume plan still requires. This table is generated from the corpus register itself and is rebuilt whenever the holdings change, so it reports the present position rather than an intention.

Veda mūla — 30 held, 30 still to source

The Saṃhitā text itself, accented. Of what is held, 26 can enter the pipeline as it stands and 4 needs recognition or conversion first.

Work Position
1 Poorvarchikam May 16Ready
101DV2021Ready
106 Aitareya Brahmana Aranyakam 2021Ready
2 Uttararchikam May16Ready
201_TS_2024_DVReady
2025__Atharva_Shounaka_SamhitaaReady
202_TB_2024_DVReady
3 Aagneyam finalReady
301 Aarchika 2016Ready
302 Prakruti Gaanam 1AReady
303 Prakruti Gaanam 2AReady
304 Chhaandogya UpanishatReady
305 Samaveda PadaPaathahReady
306 Ooha Ganam 1AReady
307 Ooha Ganam 2AReady
308 Tandya 8 BrahmanasReady
4 Aiindram finalReady
5 Paavamanam finalReady
6 Aaranyakam finalReady
Bruhadaranyaka Upanishat MoolamReady
Gopatha Brahmanam DV 2021Ready
Kaanva SamhitaaIn preparation
Kaanva Samhitaa PuBlisher FileIn preparation
Maadhyandina SamhitaaReady
Maitrayaneeya Samhita SatvalekarIn preparation
Maitrayaneeya SamhitaaReady
Samaveda Samhitaa - AarchikamReady
Taittireeya BrahmanaAranyakam AnukramanikaProof ReadingReady
Taittireeya Samhita MantraAnukramanika Proof ReadingReady

Padapāṭha — 7 held, 0 still to source

The word-by-word recitational text. Of what is held, 7 can enter the pipeline as it stands and 0 needs recognition or conversion first.

Work Position
207 Yajurveda PadaPaatha TSPDV2022Ready
305 Samaveda PadaPaathahReady
408 Atharva Padam 2025Ready
409 Gopatha Brahmanam Padapatha 2024Ready
Rigveda_Padam 1-4Ready
Rigveda_Padam 5-8Ready
TBP DV 2012Ready

Bhāṣya — 94 held, 30 still to source

Commentary — Sāyaṇa, Bhaṭṭa Bhāskara and others. Of what is held, 47 can enter the pipeline as it stands and 46 needs recognition or conversion first.

Work Position
103510524-Atharv-Ved-Part-1-Bhashya-by-Shri-Ram-Sharma-AcharyaIn preparation
103568228-Atharv-Ved-Samhita-Part-2-Bhashya-by-Shri-Ram-Sharma-AcharyaIn preparation
179354273-Shatpath-Brahman-Hindi-Vigyan-Bhashya-Dwitiya-Khanda-Motilal-Shastri-Part3In preparation
AB IndexIn preparation
Aitareya AaranyakamReady
Aitareya Brahmanam Part 1- OmkarReady
Aitareya Brahmanam Part 2- PriyadarshiniReady
Aitareya_Brahmanam_with_Sayanabhashya_Part_1_-_Kasinathsastri_Agase_1896ASS_032_In preparation
Aitareya_Brahmanam_with_Sayanabhashya_Part_2_-_Kasinathsastri_Agase_1896ASS_032_In preparation
Aitareyaranyakam_with_Sayanabhashya_-_Babasastri_Phadke_1898ASS_038_In preparation
Aranya Samhita Samaveda Sayana Bhashya Jivanand Vidya Sagar 1891 - CopyIn preparation
Arsheya Brahmana BhashyamReady
Atharva Sounaka 1-5In preparation
Atharva Sounaka 10-18In preparation
Atharva Sounaka 19-20In preparation
Atharvaveda_Bhashyam (1-5)Ready
Atharvaveda_Bhashyam (11-18)Ready
Atharvaveda_Bhashyam (19-20)Ready
Atharvaveda_Bhashyam (6-8)Ready
Atharvaveda_Bhashyam (9 -10)Ready
BBB Ashtakam 1Ready
BBB Ashtakam 2Ready
BBB Ashtakam 3Ready
BBB Kanda 1Ready
BBB Kanda 2Ready
BBB Kanda 3 (A5- 258)Ready
BBB Kanda 5 (A5- 263)Ready
BBB Kanda 6 (A5-262)Ready
BBB Kanda 7 (A5-187)Ready
BBB T AaranyakamReady
Bruhadarayaka BhashyamReady
Devatadhyaya - Samhitopanisad - Vamsa BrahmanamReady
Ganesa_Atharvasirsham_Sabhashyam_-_Vamansastri_Islampurkar_1889ASS_001_In preparation
Gopatha Meanings Old EditionReady
Praataranuvaaka Agni-Ushas-AshwinIn preparation
Rigveda Sahmita Bhashyam Ashtakam 1Ready
Rigveda Sahmita Bhashyam Ashtakam 2Ready
Rigveda Sahmita Bhashyam Ashtakam 3Ready
Rigveda Sahmita Bhashyam Ashtakam 4Ready
Rigveda Sahmita Bhashyam Ashtakam 5Ready
Rigveda Sahmita Bhashyam Ashtakam 6Ready
Rigveda Sahmita Bhashyam Ashtakam 7Ready
Rigveda Sahmita Bhashyam Ashtakam 8Ready
Rudra Adhyaaya BhashyamRetired
Rudra Adhyaaya Bhashyam 2016Ready
Rudradhyaya_with_Commentaries_of_Sayana__Bhattabhaskara_1935_ASS_002_In preparation
Saayana Bhashyam TB 2.6-3.7Ready
Saayana Bhashyam TS 1 KaandaReady
Saayana Bhashyam TS 3 KaandaReady
Saayana Bhashyam TS 4 KaandaReady
Samaveda ArsheyadeepaReady
Samaveda Sahmita BhashyamReady
Samavidhana brahmanamReady
Sandhya Vandanam With Meanings with DetailsReady
Shadvimsha brahmanamReady
TA3 Ekagni KaandamIn preparation
TAPart_1_-_Babasastri_Phadke_1898ASS_036_In preparation
TAPart_2_-_Babasastri_Phadke_1927ASS_036_In preparation
TB1.1_Part_1_-_Narayanasastri_Godbole_1934ASS_037_Ready
TB2.6_Part_2_-_Narayanasastri_Godbole_1898ASS_037_In preparation
TB3.8_Part_3_-_Narayanasastri_Godbole_1898ASS_037_In preparation
TS1.1_Part_1_-_Kasinath_Sastri_Agase_1940ASS_042_Ready
TS1.31_Part_2_-_Kasinath_Sastri_Agase_1940ASS_042_In preparation
TS1.71_Part_3_-_Kasinath_Sastri_Agase_1947ASS_042_In preparation
TS2.1_Part_4_-_Kasinath_Sastri_Agase_1946ASS_042_In preparation
TS2.51_Part_5_-_Kasinath_Sastri_Agase_1946ASS_042_In preparation
TS3.5_Part_6_-_Kasinath_Sastri_Agase_1949ASS_042_In preparation
TS5.1_Part_7_-_Kasinath_Sastri_Agase_1949ASS_042_In preparation
TS6.1_Part_8_-_Kasinath_Sastri_Agase_1951ASS_042_Ready
Taittiriyopanishat Satikaa ShaankarabhashyaIn preparation
Tandya Brahmana BhashyamReady
ekagni_kanda_haradatta_taittiriyaIn preparation
shadvimsha_brahmana with bhashyamIn preparation
shukla_yajurveda_two_commentariesIn preparation
ssk-samaveda-with-commentary-of-madhvaIn preparation
t_aranyaka_bhaskara_01In preparation
t_aranyaka_bhaskara_02In preparation
t_brahmana_bhaskara_01In preparation
t_brahmana_bhaskara_02In preparation
t_brahmana_bhaskara_03.1In preparation
t_brahmana_bhaskara_03.2In preparation
t_samhita_bhaskara_01(1.1-1.3)In preparation
t_samhita_bhaskara_02(1.4-1.6)In preparation
t_samhita_bhaskara_03(1.7-2.2)In preparation
t_samhita_bhaskara_04(2.3-2.6)In preparation
t_samhita_bhaskara_05(3.1-3.5)In preparation
t_samhita_bhaskara_06(5.1-5.4)In preparation
t_samhita_bhaskara_07(5.5-5.7In preparation
t_samhita_bhaskara_08(6.1-6.4)In preparation
t_samhita_bhaskara_09(6.5-7.3)In preparation
t_samhita_bhaskara_10(7.4-7.5)In preparation

Vedāṅga — 63 held, 21 still to source

The six auxiliary disciplines. Of what is held, 43 can enter the pipeline as it stands and 16 needs recognition or conversion first.

Work Position
01 Atharva Veda Chaturadhyayika (224 P)Ready
02 Atharva Veda Parishitham (346 P)Ready
03 Atharva Vediya Panchapatalika (26 P)Ready
04 Atharva_Praatishakhya (DV) (14 P)Ready
05 Atharva Veda Mandukeeya Shiksha (15 P)Ready
06 Atharva Veda Bhashyam (1-5 Khanda`s) - (344 P)Ready
07 Atharvaveda_Bhashyam (6-10) Kanda`s- (359 P)Ready
08 Athrvaveda_Bhashyam (11-18) Kand`s (364 P)Ready
09 Atharvaveda_Bhashyam (19-20) Kanda`s (361 P)Ready
10 Shounaka Samhitaa Padam 408DV-A4 -772PReady
11 Koushika Paddhati (Keshava Kruta) - 273PReady
11 Koushika Paddhati (Keshava Kruta) - 273P Atharva KarmaaniReady
401DV-A4. (working)Ready
Aashwalaayana Gruhya Sutram with CommentaryReady
Aranyaka Siksha 2016Ready
Aranyaka Siksha DVReady
Atharva Books List with Page NumbersReady
Atharva Veda PratishakhyamReady
Atharva_Praatishakya(TL)Ready
Atharva_Rishi_Chandas_Devata(1-20 Kand`s)Ready
Atharvaveda ChandasReady
Bharadwaja ShikshaReady
Jata DarpanamReady
Kaala Nirnaya PattikaaReady
Lakshana Grantha RigvedaReady
Lakshana Moolam AtharvaReady
Lakshana Moolam YajurvedaReady
Manduki ShikshaIn preparation
Naradiya Shiksha Commentary 2In preparation
Naradiya Shiksha with Bhatta Shobhakar's Shiksha Vivarana Commentary - NaradIn preparation
Others_Yohi -Prapti with commentaryIn preparation
Others_yohi_prapti_shikshaIn preparation
Panchavidha SutramReady
Pushpa Sutram of Samaveda_5262__Alm_24_Shlf_1_Devanagari - Sutra PaddhatiIn preparation
PushpasutramReady
Pushpasutram MoolamReady
Rigveda Brahmakarma samuchayaIn preparation
Rigveda Praatishaakhyam with Uvata BhaashyamReady
Rigveda Pratishaakhyam - Prayaga PrintIn preparation
Rik Pratishakhya Moolam 2023Ready
Rik TantramReady
Saama TantramReady
Sama Lakshana StabakamReady
Sama ModelReady
Sapta Lakshanam - PriyadarshiniReady
Sapta Lakshanam 1st EditionIn preparation
Shiksha Yajur Moolam 2Ready
Shukla Yajurveda PratishakhyaReady
Taittireeya PraatishaakhyamDVIn preparation
Taittireeya Pratishaakhyam 2 VyaakhyaIn preparation
Taittireeya Pratishakhayam SavyakhyamReady
Taittireeya Pratishakhyam Brief EnIn preparation
Taittiriya Praatishaakhyam 2022 With 3 CommentariesReady
Taittiriya-Pratisakhya MahisheyaIn preparation
Taittiriya-Pratisakhya WhitneyIn preparation
Vyaasa Shiksha VedaTaijasa Sarvalakshana ManjariIn preparation
Vyasa Shiksha Vyakhya 2020Ready
Yohi Shiksha - PriyadarshiniReady
sama_veda_pratishakhyaIn preparation

Śrauta prayoga — 29 held, 26 still to source

Ritual manuals and their sequence. Of what is held, 29 can enter the pipeline as it stands and 0 needs recognition or conversion first.

Work Position
20260108_000932_bhagpur-01Ready
20260120_140835_Agnihotra PrayogaReady
20260122_222646_Yajusha Shraaddha PrayogahReady
20260131_084705_Agnishtoma_1Ready
20260131_084725_Agnishtoma_2Ready
20260131_084732_Agnishtoma_3Ready
20260131_085930_Agnishtoma_5Ready
20260131_090831_Agnishtoma_41Ready
20260131_090839_Agnishtoma_42Ready
20260131_091027_chaturmasya1Ready
20260131_091053_chaturmasya2Ready
20260201_155430_Ramayana MuktaavaliReady
20260202_113118_saraswati_vidya_prarthanamReady
20260210_152838_Sachchidananda Neeti Maala 2018 FinalReady
20260217_230159_Apara Prayoga (Bharatula)Ready
20260221_054507_RUDRA PRAPANC FINAL BOOKReady
20260309_234910_plan1Ready
20260312_024159_sample_nirnaya_sagarReady
20260421_233658_154075388-Asvalayana-Srautasutra-1917-pdf_compressedReady
20260525_012330_Oudgaatra_AgnishtotmaReady
20260526_023432_MA_SanskritReady
20260617_233440_missing pages PravargyaReady
20260617_234656_1Ready
20260625_085924_Agnyadheeya prayogaReady
20260625_090425_AnvarambhaneeyaReady
20260625_091149_Niroodha pashubandha prayogaReady
20260625_092223_Chaturmaasya prayoga 2Ready
20260625_092814_Chaturmaasya prayogaReady
20260625_093752_Niroodha pashubandha prayogaReady

Śāstra — 4 held, 75 still to source

Mīmāṃsā, Nyāya, Vedānta, Vyākaraṇa. Of what is held, 4 can enter the pipeline as it stands and 0 needs recognition or conversion first.

Work Position
20260109_003401_Mimamsa Nyaya Prakasha NSP 1_textReady
20260120_130641_Pratibandhakata Vada Gadhadhara Narayana Shastri PatwardhanReady
20260309_234940_Ananda Giri TeekaReady
20260309_235044_Vedanta Sutra MuktavaliReady

Lexical sources — 63 held, 0 still to source

The kośa collections held for corroboration. Of what is held, 63 can enter the pipeline as it stands and 0 needs recognition or conversion first.

Work Position
AbhidhanacintamaniReady
Abhidhanacintamani (Hemacandra)Ready
Abhidhanacintamani - ParisistaReady
Abhidhanacintamani - SilonchaReady
AbhidhanaratnamalaReady
Abhidhanaratnamala (Halayudha)Ready
Amarakosa (ontology build)Ready
Amarakosa with Sudha commentaryReady
AnekarthadhvanimanjariReady
Apte, English-Sanskrit DictionaryReady
Apte, The Practical Sanskrit-English DictionaryReady
Benfey, Sanskrit-English DictionaryReady
Boehtlingk & Roth, Sanskrit-Woerterbuch (7 Baende)Ready
Boehtlingk, Sanskrit-Woerterbuch in kuerzerer FassungReady
Bopp, Glossarium SanscritumReady
Borooah, English-Sanskrit DictionaryReady
Burnouf, Dictionnaire classique Sanscrit-FrancaisReady
Cappeller, Sanskrit-English DictionaryReady
Cappeller, Sanskrit-WoerterbuchReady
EkaksaranamamalaReady
Goldstuecker, Sanskrit-English DictionaryReady
Grassmann, Woerterbuch zum Rig-VedaReady
L. R. Vaidya, Sanskrit-English DictionaryReady
Lanman, Sanskrit Reader vocabularyReady
Macdonell, A Practical Sanskrit DictionaryReady
Monier-Williams (1872 edition)Ready
Monier-Williams, A Sanskrit-English DictionaryReady
Monier-Williams, English-Sanskrit DictionaryReady
Sabda-Sagara, Sanskrit-English DictionaryReady
Sabdakalpadruma (Radhakantadeva)Ready
Sabdakalpadruma (StarDict build)Ready
Soerensen, Index to the Names in the MahabharataReady
Vacaspatyam (StarDict build)Ready
Vacaspatyam (Taranatha Tarkavacaspati)Ready
Vedic Index of Names and Subjects (Macdonell & Keith)Ready
Wilson, Sanskrit-English DictionaryReady
Yates, Sanskrit-English DictionaryReady
pwkvnReady

Derivation sources — 5 held, 0 still to source

Inputs to the Pāṇinian generator. Of what is held, 5 can enter the pipeline as it stands and 0 needs recognition or conversion first.

Held items in this division are lexical sources whose individual titles are not published.

Kāvya, Itihāsa, Purāṇa — 0 held, 60 still to source

The surrounding literature. Of what is held, 0 can enter the pipeline as it stands and 0 needs recognition or conversion first.

Miscellany — 4 held, 61 still to source

Everything not yet classified. Of what is held, 0 can enter the pipeline as it stands and 4 needs recognition or conversion first.

Held items in this division are lexical sources whose individual titles are not published.
What is not listed. 32 held works are counted in the totals above but not named — either they are third-party lexica held for scholarly comparison rather than redistribution, or their licence has not yet been verified. A work is named here only on affirmative clearance, never by default. Works still being catalogued are omitted until their identification is settled.
Section 6

Pending uploads

Every held artefact is classified by what stands between it and the pipeline. Nothing has been loaded to the production corpus yet — the schema is still under scholarly review — so the whole holding is queued.

State Meaning Files Size
Ready to loadText-bearing files that need no conversion.156569 MB
Awaiting assemblyRecognition already complete; output not yet assembled.301,300 MB
Awaiting triageTo be checked for a text layer before any recognition.341,085 MB
Awaiting conversionLegacy word-processor formats.27 MB
SupersededA better copy of the same work is held.510 MB
The immediate opportunity. A substantial block of text recognition has already been completed and paid for, but the output was never assembled into documents. Assembling it requires no new expenditure and is the largest body of readable text available to the project today. It is unit SVU-027 in the plan below.
Section 7

Acquisition gaps

The seventy-two volume plan names 390 distinct works. Reconciling that requirement against the holdings shows precisely what has still to be sourced.

Position Distinct works
Held19
Held, pending verification68
Not yet acquired303

Where the gaps concentrate

Category Works to acquire
Darśana / Philosophy34
Saṃhitā30
Saṅgīta / Nāṭya23
Purāṇa21
Dharmasūtra / Smṛti19
Kāvya / Alaṃkāra18
Brāhmaṇa15
Stotra / Nāmāvali14
Upaniṣad14
Gṛhyasūtra13

The holdings are strongest in Veda mūla, bhāṣya and vedāṅga — the Foundation's own scholarly territory — and thinnest in the darśana, purāṇa, kāvya and applied-śāstra divisions. Acquisition is therefore sequenced by what the volumes actually require, not by what is easiest to obtain.

Section 8

Project tasks

Four stages, 160 units. Stage numbering is thematic; several stages run concurrently. Items marked Awaiting decision are held pending a scholarly determination and are deliberately not started, because the determination may change the work.

Stage 0 · Foundation & Governance — 14 units, 6 complete

Infrastructure, schema, governance and the registers that everything else is tracked against.

Unit Task Track Status
SVU-001Approve the corpus database design and create itGovernanceAwaiting decision
SVU-002Add the dictionary tables to the corpus databaseGovernanceAwaiting decision
SVU-003Commission the two computing serversInfraComplete
SVU-004Install the databases and search services on the serversInfraComplete
SVU-005Constitute the scholarly review panelGovernancePlanned
SVU-006Declare the licence for each of our three own worksGovernanceAwaiting decision
SVU-007Repository governance and access controlGovernancePlanned
SVU-008Number every document and keep a register of themGovernanceComplete
SVU-009Register every text we hold, with its loading statusCorpusComplete
SVU-010Match the books the 72 volumes need against what we holdCorpusComplete
SVU-011Generate project documents reproducibly, and reject invalid filesToolingComplete
SVU-012Corpus resilience and off-site replicationRiskPlanned
SVU-013Define what each group hands to the next, and whenGovernanceReady to start
SVU-014Bring in outside funding through non-profits and matching grantsGovernancePlanned

Stage 1 · Collect & Load the Corpus — 30 units, 6 complete

Encoding the accent correctly, then loading the text — mūla, pada-pāṭha, bhāṣya and the ancillary śāstras — with provenance intact.

Unit Task Track Status
SVU-020Count every special VijayaDV character in the corpusEncodingComplete
SVU-021Sort the VijayaDV special characters into accent, marker and sandhiEncodingComplete
SVU-022Decide the standard character each remaining accent mark maps toEncodingAwaiting decision
SVU-023Confirm that position-variant accent glyphs mean one accentEncodingAwaiting decision
SVU-024Convert the mūla of all four Vedas to standard charactersEncodingAwaiting decision
SVU-025Prove every converted file converts back unchangedEncodingAwaiting decision
SVU-026Load the 120 files that are already machine-readableLoadingAwaiting decision
SVU-027Assemble the volumes already put through OCRLoadingComplete
SVU-028Check each un-OCR'd PDF for existing text before paying for OCRLoadingComplete
SVU-029Run OCR on the books that are genuine scansLoadingAwaiting decision
SVU-030Extract the text from PDFs that already carry itLoadingReady to start
SVU-031Convert the two old-format Kāṇva Saṃhitā filesLoadingComplete
SVU-032Archive the 5 PDFs whose DOCX we already holdLoadingIn progress
SVU-033Load the Kalpa Sūtra collection, split at sūtra levelCorpusAwaiting decision
SVU-034List every work quoted in the 72 volumesCorpusReady to start
SVU-035Put the 303 missing works in the order we should obtain themCorpusComplete
SVU-036Bring the outside digital corpora into our own formatCorpusPlanned
SVU-037Mark where each ṛk, sūtra and śloka begins and endsCorpusPlanned
SVU-038Freeze the numbering that addresses every sentence in the corpusArchitectureAwaiting decision
SVU-039Index every passage for meaning-based searchRetrievalPlanned
SVU-040Link ṛk to devatā, sūkta to ṛṣi, mantra to riteRetrievalPlanned
SVU-041Load the four 2020 Bhāṣya Pilot volumes as the gold standardMAPAwaiting decision
SVU-042Resolve every grammatical and śāstric citation in the BhāṣyaMAPPlanned
SVU-043Group each mantra and brāhmaṇa passage with its Padapāṭha and BhāṣyamMAPPlanned
SVU-044Compute meaning, analysis and presentation for a mantra on demandMAPPlanned
SVU-044.1Spec — MAP analysis contract — what the reader is shown and from whatMAPReady to start
SVU-044.2Build — Compute meaning, analysis and presentation for a mantra on demandMAPPlanned
SVU-046Publish finished material to vaakya.vedanidhi.in as it is readyDeliveryPlanned
SVU-048Reconcile filing of recently added documentsHygieneIn progress
SVU-049Recover the text from PDFs written in legacy fontsLoadingReady to start

Stage 2 · Pada Kośa — 38 units, 3 complete

The lexical foundation: the conventional vocabulary drawn from the kośas, the Vedic vocabulary drawn from pada-pāṭha and bhāṣya, and the compositional vocabulary generated from Pāṇini's rules.

Unit Task Track Status
SVU-050Load the 62 dictionaries and make their words searchablePada Kośa AAwaiting decision
SVU-051Publish the census of what the dictionaries containPada Kośa AComplete
SVU-052Access and serving policy for third-party lexical sourcesPada Kośa AAwaiting decision
SVU-053Fix which dictionary the reader is shown firstPada Kośa AAwaiting decision
SVU-054Flag the 106,014 words attested in only one dictionaryPada Kośa AAwaiting decision
SVU-055Load every Padapāṭha we hold, with its accents intactPada Kośa VAwaiting decision
SVU-057Generate a Padapāṭha for texts that lack onePada Kośa VPlanned
SVU-057.1Spec — Rules for generating Pada-pāṭha where none is attestedPada Kośa VReady to start
SVU-057.2Build — Generate a Padapāṭha for texts that lack onePada Kośa VPlanned
SVU-058Build the Vedic lexicon: words and meanings drawn from the BhāṣyaPada Kośa VPlanned
SVU-058.1Spec — Method for extracting the Vaidika lexicon from BhāṣyaPada Kośa VReady to start
SVU-058.2Build — Build the Vedic lexicon: words and meanings drawn from the BhāṣyaPada Kośa VPlanned
SVU-059Derive words from dhātu and pratyaya, with the sūtra chainPada Kośa BComplete
SVU-060Fix how accents are printed in published formsPada Kośa BAwaiting decision
SVU-061Decide whether pracaya and ekaśruti are appliedPada Kośa BAwaiting decision
SVU-062Measure our derived accents against the gold corpus, vowel by vowelPada Kośa BAwaiting decision
SVU-062.1Spec — Accent-agreement test design and acceptance thresholdPada Kośa BAwaiting decision
SVU-062.2Build — Measure our derived accents against the gold corpus, vowel by vowelPada Kośa BPlanned
SVU-063Decide how many derived forms we generatePada Kośa BAwaiting decision
SVU-064Generate the full set of forms at the agreed sizePada Kośa BAwaiting decision
SVU-065Build the fast word-lookup indexPada Kośa BAwaiting decision
SVU-066Write the ~3,400 morpheme meanings in Telugu, Kannada, Hindi and EnglishPada Kośa BPlanned
SVU-067Give the nirvacana of a word, with competing etymologiesNiruktaAwaiting decision
SVU-067.1Spec — Nirvacana model — sources, competing etymologies, presentationNiruktaReady to start
SVU-067.2Build — Give the nirvacana of a word, with competing etymologiesNiruktaPlanned
SVU-068Establish the liṅga of each stem from the liṅga authoritiesVyākaraṇaPlanned
SVU-069Apply sandhi by rule, with the accent change at each juncturaPāṭhaPlanned
SVU-069.1Spec — Sandhi rule inventory and svara behaviour at each juncturaPāṭhaReady to start
SVU-069.2Build — Apply sandhi by rule, with the accent change at each juncturaPāṭhaPlanned
SVU-070Generate the Krama pāṭha of any Saṃhitā passagePāṭhaPlanned
SVU-070.1Spec — Krama construction rulesPāṭhaReady to start
SVU-070.2Build — Generate the Krama pāṭha of any Saṃhitā passagePāṭhaPlanned
SVU-071Generate the eight Vikṛti pāṭhas, with accents preservedPāṭhaPlanned
SVU-071.1Spec — The eight Vikṛti forms — construction rules per formPāṭhaReady to start
SVU-071.2Build — Generate the eight Vikṛti pāṭhas, with accents preservedPāṭhaPlanned
SVU-072Varṇa Krama — already shipped, so settle reuse terms onlyPāṭhaComplete
SVU-073Look up a form and return its analyses for the hoverPada KośaIn progress
SVU-074Join a Veda word to its dictionary entry and its derivationPada Kośa VAwaiting decision

Stage 3 · Synthesis, Models & Applications — 60 units, 0 complete

Cross-referencing, translation, the generative engines, and the reading application built on top of them.

Unit Task Track Status
SVU-080Connect cross-references across Veda, Vedāṅga, Upāṅga and UpavedaSynthesisPlanned
SVU-081Translate mantras into Indian languages, grounded in the BhāṣyaSynthesisPlanned
SVU-081.1Spec — Anuvāda method — grounding every rendering in BhāṣyaSynthesisReady to start
SVU-081.2Build — Translate mantras into Indian languages, grounded in the BhāṣyaSynthesisPlanned
SVU-082Find passages similar in word and in meaningSynthesisPlanned
SVU-082.1Spec — Samānatā — what counts as similarity in śabda and in arthaSynthesisReady to start
SVU-082.2Build — Find passages similar in word and in meaningSynthesisPlanned
SVU-083Analyse a sentence across all the verticalsSynthesisPlanned
SVU-084Assemble the instruction-and-answer set that trains the modelModelPlanned
SVU-085Fine-tune the base model on the Vedic corpusModelPlanned
SVU-086Answer from retrieved passages rather than from memoryModelPlanned
SVU-087Improve answers from scholars' ratings of paired outputsModelPlanned
SVU-088Re-measure training time on the servers we actually haveModelAwaiting decision
SVU-089Generate a complete Śrauta rite sequence, every step citedEnginePlanned
SVU-089.1Spec — Śrauta Prayoga sequence model and citation requirementsEngineReady to start
SVU-089.2Build — Generate a complete Śrauta rite sequence, every step citedEnginePlanned
SVU-090Generate Sāmaveda stotras with viṣṭuti and stobhaEnginePlanned
SVU-090.1Spec — Stotriyā assembly, viṣṭuti patterns and stobha rulesEngineReady to start
SVU-090.2Build — Generate Sāmaveda stotras with viṣṭuti and stobhaEnginePlanned
SVU-091Teach the model the Vikṛti pāṭhas from rule-generated formsEnginePlanned
SVU-092Analyse arguments in Tarka form: pañcāvayava and Navya-NyāyaEnginePlanned
SVU-092.1Spec — Śābda-bodha and pariṣkāra representation for TarkaEngineReady to start
SVU-092.2Build — Analyse arguments in Tarka form: pañcāvayava and Navya-NyāyaEnginePlanned
SVU-094Analyse a recitation's sound: svara, modulation and volumeAudioPlanned
SVU-094.1Spec — Svara extraction and notation conventions for SāmagānaAudioReady to start
SVU-094.2Build — Analyse a recitation's sound: svara, modulation and volumeAudioPlanned
SVU-095Re-frame a melody onto a new chandasEnginePlanned
SVU-095.1Spec — Ūha — how melody is re-framed onto new chandasEngineReady to start
SVU-095.2Build — Re-frame a melody onto a new chandasEnginePlanned
SVU-097Interpret a passage through four darśanas side by sideEnginePlanned
SVU-097.1Spec — The four lenses — scope and retrieval boundary of eachEngineReady to start
SVU-097.2Build — Interpret a passage through four darśanas side by sideEnginePlanned
SVU-098Show a passage in five layers, mūla to related passagesAppIn progress
SVU-099Show every word's grammar on hover, by lookup not guessworkAppIn progress
SVU-100Rite generator for scholars, with its citation trailAppPlanned
SVU-101Screen that turns a Saṃhitā passage into all eight Vikṛti formsAppPlanned
SVU-102Sāma Studio: stotra generator, notation viewer, sound panelAppPlanned
SVU-103Tarka workbench, and the akṣara-by-akṣara Varṇa Krama viewAppPlanned
SVU-104Stotra library by devatā and chandas, and an etymology explorerAppPlanned
SVU-105Notes connecting a passage to science, philosophy and the artsAppPlanned
SVU-106Serve the model on the second node, with cachingInfraPlanned
SVU-107Run the pilot with 25–50 faculty and scholarsValidationPlanned
SVU-108Score answers on all 2,000 Swadharma topics against scholar answersValidationPlanned
SVU-109Check generated rites against the 2026 Agniṣṭoma performanceValidationPlanned
SVU-160Improve the speech-to-text model on Vedic recitationAudioIn progress
SVU-161Record volunteers chanting each śākhā, with correct intonationAudioReady to start
SVU-162Speak Vedic text aloud with the svara correctAudioPlanned
SVU-163Settle the licence and access terms for the training audioAudioReady to start
SVU-170Test every group's output for function and for contentTestingReady to start
SVU-171Sample scholarly quality: random at first, then systematicTestingReady to start
SVU-175Final review before anything reaches the publicReviewPlanned
SVU-140Approve the Mīmāṃsā sentence-classification vocabularies and record formatsMVVFIn progress
SVU-141Tag the Darśapūrṇamāsa pilot passage by handMVVFPlanned
SVU-142Tag sentences automatically from the visible marks in the textMVVFPlanned
SVU-143Review console where a scholar confirms or overrides each automatic tagMVVFPlanned
SVU-144Enter the first Adhyāya of the Nyāyamālā as Adhikaraṇa recordsMVVFPlanned
SVU-145Enter all the Nyāyamālā Adhikaraṇas, with Bhāṭṭa and Prābhākara positionsMVVFPlanned
SVU-146Registry of rites and episodes, so one sentence can be traced through every riteMVVFPlanned
SVU-147Three services: classify a sentence, explain its tag, walk the episode graphMVVFPlanned
SVU-148Training set of tag decisions with the reasoning behind eachMVVFPlanned

Stage 4 · Presentation & Outreach — 18 units, 0 complete

Publication, the researcher interface, and bringing the material to a general audience.

Unit Task Track Status
SVU-115Put this plan and its current status on sarvaveda.infoSiteReady to start
SVU-116Publish pages as fixed HTML with sitemaps, so search engines read themSitePlanned
SVU-117Let a reader's correction on the page become a proposed changeSiteAwaiting decision
SVU-118Open the site with General, Scholar and Institutional accessReleasePlanned
SVU-119Machine access for researchers, documented and rate-limitedReleasePlanned
SVU-121Answer any topic from the śāstra, with sourced citationsReachPlanned
SVU-122Production route from corpus to published video, audio and webMediaPlanned
SVU-123Run regular scholarly webinars from the corpusMediaPlanned
SVU-124Generate the VVS textbook series from the validated corpusMediaPlanned
SVU-125Produce short videos at scale, each scholar-approved before releaseMediaPlanned
SVU-180Presentation that explains the initiative to a general audienceMediaReady to start
SVU-126Publish daily across the channels, including on current eventsMediaPlanned
SVU-128Publish the consultation papers with their licence settledReleaseAwaiting decision
SVU-129Retrain each quarter on new volumes and feedbackReleasePlanned
SVU-130Transcribe 1,000 hours of video and align it to the textCorpusPlanned
SVU-149Public Mīmāṃsā pages: corpus browser, tag explorer, Adhikaraṇa readerMVVFPlanned
SVU-150Reorganise vedavishtaram.in into seven top-level sectionsSiteReady to start
SVU-151Build the VVS catalogue straight from the Master RegisterSitePlanned
How to read this. A unit is complete only when its definition of done is met and a test asserts it. "Ready to start" means no decision and no prerequisite stands in the way. The plan is a living document and unit numbers are permanent, so this page can be compared honestly against itself over time.