SarvaVeda is a sovereign foundation LLM and research platform built on the 72-volume Veda Vijñāna Viṣṭaram Series, the complete Kalpa Sūtra literature, all eighteen Mahāpurāṇas, and the full Ṣaḍaṅga corpus — self-hosted, scholar-validated, and answerable to the tradition itself.
Every text enters the model through a graded trust system. The higher the stratum, the greater its weight in training — and every sentence remains traceable to the printed page it came from.
The 72 volumes of the Veda Vijñāna Viṣṭaram Series — 15,000 pages of curated exposition spanning Veda, Vedāṅga, Itihāsa, Purāṇa, Āgama, Jyotiṣa, Sacred Geography, and the musical traditions of Bhārata. The supervised training spine of the model. Every volume, with its production status →
The 2,000-topic Swadharma Master List; the SGSP Kalpa Sūtra Unicode collection — eleven Śrauta Sūtras, the Gṛhya register, commentated Dharma and Śulba texts — and curated assessment banks from SVAMI.
Hundreds of scanned editions passing through the seven-stage industrial OCR pipeline: critical editions, traditional bhāṣyas of Sāyaṇa and the commentators, all eighteen Mahāpurāṇas, the Stotra literature, and the complete Kalpa Sūtra library in four classes.
1,000 hours of transcribed lectures by traditional scholars, aligned with the text corpus — carrying the oral tradition's explanatory voice into the model.
The four strata above describe the corpus in principle. For what is actually in it — every text, its source edition, its trust tier and its verse count, regenerated whenever a module completes — see index.sarvaveda.info, the Master Index.
All eighteen Mahāpurāṇas are in. Every figure below is counted from the corpus database, not entered by hand, and is regenerated whenever a module completes — Vedāṅga and Itihāsa and Purāṇa so far, 2 of the 11 departments.
A verse count marked * includes verses carrying a known doubt — 23 of the 43 texts. The doubt is named on the row: a generated script that will not convert back, a character no romanised Sanskrit should hold, a verse numbered by position because the source printed no number, or a locus the source repeats. The verses are all present and addressable; the asterisk says only that a reader should not treat the figure as settled.
BaudhSSsilver12,407SGSP Kalpa Sūtra Unicode collectionApSSsilver8,331SGSP Kalpa Sūtra Unicode collectionSankhSSsilver6,368SGSP Kalpa Sūtra Unicode collectionBharSSsilver919SGSP Kalpa Sūtra Unicode collectionManSSsilver802SGSP Kalpa Sūtra Unicode collectionVarSSsilver557SGSP Kalpa Sūtra Unicode collectionMasSSsilver310SGSP Kalpa Sūtra Unicode collectionKatySSsilver256SGSP Kalpa Sūtra Unicode collectionAsvSSsilver212SGSP Kalpa Sūtra Unicode collectionVaitSSsilver136SGSP Kalpa Sūtra Unicode collectionJaimSSsilver28SGSP Kalpa Sūtra Unicode collectionBaudhGSsilver3,303SGSP Kalpa Sūtra Unicode collectionHirGSsilver1,193SGSP Kalpa Sūtra Unicode collectionParGSsilver813SGSP Kalpa Sūtra Unicode collectionAsvGSsilver762SGSP Kalpa Sūtra Unicode collectionVaikhGSsilver426SGSP Kalpa Sūtra Unicode collectionApGSsilver409SGSP Kalpa Sūtra Unicode collectionKhadGSsilver21SGSP Kalpa Sūtra Unicode collectionMBhsilver73,816BORI critical edition (Bhandarkar Oriental Research Institute, Pune) — John Smith's revision of M. TokunagaRambronze18,761*Baroda critical edition text (M. Tokunaga, revised by John Smith), via GRETIL · * 32 whose generated script does not convert back · Kāvya by received classification; filed here with the Itihāsa on VKG's rulingPadPreference49,688*Saṃskṛta Wikisource dump, mainspace · * 269 whose generated script does not convert back; 666 numbered by position because the source printed none; 301 whose locus repeats in the source, filed as ~2BvPreference20,757*Saṃskṛta Wikisource dump, mainspace · * 982 whose generated script does not convert back; 21 numbered by position because the source printed none; 640 whose locus repeats in the source, filed as ~2BhavPreference18,629*Saṃskṛta Wikisource dump, mainspace · * 1,278 whose generated script does not convert back; 276 numbered by position because the source printed none; 208 whose locus repeats in the source, filed as ~2NarPbronze15,586*Sansknet project · * 40 whose generated script does not convert back; 5 carrying characters no romanised Sanskrit should hold; 20 whose locus repeats in the source, filed as ~2 · holds Parts 1 and 2BhPbronze14,062*GRETIL per-skandha files · * 33 whose generated script does not convert back; 9 whose locus repeats in the source, filed as ~2 · holds Skandhas 1-12BndPbronze13,743*Bombay: Veṅkaṭeśvara Steam Press · * 4 whose generated script does not convert back; 3 carrying characters no romanised Sanskrit should hold; 36 whose locus repeats in the source, filed as ~2 · holds 3 parts; verses 2,70.3-49 not available at presentBrPbronze13,579*Tübingen Purāṇa Project (Schreiner / Söhnen-Thieme) · * 19 whose generated script does not convert back; 2 carrying characters no romanised Sanskrit should hold; 1 whose locus repeats in the source, filed as ~2 · holds Adhyāyas 1-246GarPbronze11,963*Bombay: Veṅkaṭeśvara Steam Press · * 3 whose generated script does not convert back; 1 carrying characters no romanised Sanskrit should hold; 10 whose locus repeats in the source, filed as ~2 · holds Parts 1-3APbronze11,387*Bibliotheca Indica 65,1-3 · * 40 whose generated script does not convert back; 14 carrying characters no romanised Sanskrit should hold; 64 whose locus repeats in the source, filed as ~2VarPreference10,428*Saṃskṛta Wikisource dump, mainspace · * 225 whose generated script does not convert back; 5 numbered by position because the source printed none; 385 whose locus repeats in the source, filed as ~2LiPbronze9,135*Bombay: Veṅkaṭeśvara Steam Press 1906 · * 5 whose generated script does not convert back; 2 carrying characters no romanised Sanskrit should hold · holds Part 1 adhy. 1-108; part 2 adhy. 1-55MatsPbronze8,498*Calcutta: Caukhamba Vidyabhavan 1954 · * 3 whose generated script does not convert back; 2 carrying characters no romanised Sanskrit should hold · holds Adhyāyas 1-176KūrmPbronze5,823*Sansknet project · * 1 whose locus repeats in the source, filed as ~2 · holds Parts 1 and 2VamPbronze5,683*A.S. Gupta (ed.), Varanasi: All India Kashiraj Trust 1967 · * 6 whose generated script does not convert back; 2 carrying characters no romanised Sanskrit should hold · holds Adhyāyas 1-69, plus the Saromāhātmya inserted after adhy. 23ŚivPbronze5,667*Bombay: Veṅkaṭeśvara Steam Press c. 1920 (input Jun Takashima) · * 15 carrying characters no romanised Sanskrit should hold · holds Book 1 Vidyeśvara-saṃhitā; book 7 Vāyavīya-saṃhitā parts 1-2ViPbronze5,178*Critical edition, M.M. Pathak, Vadodara: Oriental Institute 1997-1999 · * 5 whose generated script does not convert back; 2 carrying characters no romanised Sanskrit should hold; 38 whose locus repeats in the source, filed as ~2MarkPbronze4,525*Sansknet project · * 7 carrying characters no romanised Sanskrit should hold · holds Adhyāyas 1-93 (94ff. not available at present)SkPbronze1,647*Critical edition, Adriaensen / Bakker / Isaacson, Groningen: Egbert Forsten 1998- · * 1 whose generated script does not convert back · holds Adhyāyas 1-31.14 (to be continued)NsPbronze3,321*Siddheswar Jena (ed.), Delhi: Nag 1987 · * 8 whose generated script does not convert back; 78 carrying characters no romanised Sanskrit should holdDgbronze510input Ursula Honegger · holds Devībhāgavata Purāṇa 7,31-40 onlyVdhabronze303*Bombay 1912 · * 2 whose generated script does not convert back; 2 carrying characters no romanised Sanskrit should hold; 2 whose locus repeats in the source, filed as ~2 · holds Adhyāyas 3,343-353 and 2,127 onlyRKVbronze7,776*Kṣemrāj Śrīkṛṣṇadās (ed.), Bombay 1910 · * 4 whose generated script does not convert back · holds Revākhaṇḍa — the only Vāyu material GRETIL holdsRKSbronze6,721*Oṅkārānanda Giri (ed.), Hośaṅgābāda 1994 · * 1 whose generated script does not convert back · holds Revākhaṇḍa, completeLast module loaded 2026-08-22. Every verse is addressable by its own citation — MBh.1.1.1, PadP.6.253.1 — and verifies byte-identical against the edition it came from. Module status, declared-against-ingested counts and every sub-project surface are at index.sarvaveda.info.
A 70-billion-parameter foundation model, fine-tuned on the corpus and extended by nine śāstra-specific engines — each one auditable, each one citing its sources.
Complete rite sequences — ṛtvij assignments, mantra placement with svara, havirdravya, kāla-nirnaya — every step carrying a traceable sūtra citation.
Stotriyā assembly, viṣṭuti patterns and stobha insertion for any Soma rite, in traditional numeric-svara and sargam notation.
All eight Vikṛti forms — Jaṭā to Ghana — for any Saṃhitā segment, with rule-verified sandhi and svara transformation at every junctura.
Pañcāvayava syllogisms, pūrvapakṣa–siddhānta frames, and Navya-Nyāya technical analysis of any śāstric argument.
Akṣara-by-akṣara decomposition with svara and mātrā, built directly on Śikṣā and Prātiśākhya rules — pāṭhaśālā-grade output.
Audio in, notation out: pitch and svara extraction from recorded Sāmagāna, automatic notation, and the Ūha engine that re-frames melody onto new chandas.
The full Pāṇinian word-space — over 100 million forms, each with its complete sūtra-by-sūtra derivation — compiled into an instant-lookup lexicon.
Traditional nirvacana for any word, with competing etymologies cited, alongside the formal Pāṇinian vyutpatti — two tracks, one panel.
The same sentence interpreted in parallel through Tarka, Vyākaraṇa, Mīmāṃsā and Vedānta — the śāstrārtha assembly, computed.
In the traditional assembly, a naiyāyika, a vaiyākaraṇa, a mīmāṃsaka and a vedāntin each examine the same vākya with different instruments. SarvaVeda computes all four readings in parallel. Select a lens:
The sentence encodes a means–end inference: yāga is established as the sādhana for the sādhya, svarga. Reconstructed as a formal pañcāvayava: pratijñā — the desirer of heaven should perform yāga; hetu — because yāga is the instrument of heaven; udāharaṇa — whatever is an instrument to a desired end is to be undertaken by one who desires that end.
The analysis then tests the hetu for hetvābhāsa and states the cognition in Navya-Nyāya terms: svarga-niṣṭha-sādhyatā-nirūpita-sādhanatā resides in yāga.
Illustrative output — the production engine cites its sources for every claim.
Hundreds of scanned books — critical editions, commentaries, Grantha and Telugu texts — become verified, svara-correct digital corpus at a rhythm of 15–25 books a month.
Two details set this pipeline apart. Vedic accent marks — which ordinary OCR silently discards — are recovered graphically or transferred from verified sources. And every OCR'd word is checked against the Vyākaraṇa engine's 100-million-form lexicon: the most complete Sanskrit spell-checker ever constructed, born from the same Pāṇinian generator that powers the grammar display.
Two GPU servers in operation on SGSP-controlled hardware, with vector and graph databases deployed and the scholar review panel constituted.
The 72 volumes, the Kalpa Sūtra Unicode collection, the Purāṇa and Stotra corpora ingested; the industrial OCR programme launched; the Vyākaraṇa lexicon generated; the citation register of every text quoted in the Series compiled.
QLoRA fine-tuning of the 70B base model on the Gold and Silver corpora; retrieval-augmented generation live; scholar-graded evaluation and feedback training.
Prayoga, Stotra, Vikṛti, Tarka, Varṇa Krama, Nirukta and the Chatur-Darśana engines trained; the Sāma sound-analysis model delivers automatic notation; the web application opens for internal scholarly use.
SVAMI faculty and invited scholars test against the 2,000-topic benchmark; Prayoga outputs validated against the 2026 Agniṣṭoma performance records; the Ūha engine faces its Sāmaveda examiners.
Tiered public access — general, scholar, institutional — with the diaspora network of paṭhaśālās and centres as first subscribers, and a continuous-learning cycle absorbing each new volume and lecture.
SarvaVeda is not a standalone software project. It is the digital arm of a working Chaturveda Vidyālaya that has taught all four Vedas in the traditional gurukula manner since 2015 — the schools, faculty, admissions and projects of Veda Vijñāna Viṣṭaram are set out in full here.
Sanatana Guru Sampradaya Pratishthanam, Mysore — the umbrella trust, and custodian of the VVV Series the corpus draws on.
The teaching and publishing body whose 72-volume Series forms the model's training spine — twelve schools, nine faculty, all four Vedas with their śākhās and Ghana Pāṭha. The Series → · The institution →
The academy whose faculty serve as the scholar review panel — validating corpus quality and grading every model output.
USA 501(c)(3) — the vehicle for tax-deductible international support and diaspora institutional partnerships.
SarvaVeda is seeking founding supporters, a Technical Director (AI/ML), and institutional partners. US donations are tax-deductible through Dharma Poshanam Inc. (501(c)(3)). Scholars, engineers, and patrons who wish to be part of this work are invited to write to the Pratishthanam.
Write to the project