A Project of VKG Foundation · Sanatana Guru Sampradaya Pratishthanam · DharmaPoshanam

The complete Vedic textual universe, encoded as living intelligence.

SarvaVeda is a sovereign foundation LLM and research platform built on the 72-volume Veda Vijñāna Viṣṭaram Series, the complete Kalpa Sūtra literature, all eighteen Mahāpurāṇas, and the full Ṣaḍaṅga corpus — self-hosted, scholar-validated, and answerable to the tradition itself.

अनन्ता वै वेदाःThe Vedas are indeed infinite — Taittirīya tradition
SarvaVeda LLM — a flourishing tree whose four branches carry the Ṛgveda, Yajurveda, Sāmaveda and Atharvaveda, rooted in Om
72
VVV Volumes · Gold Corpus
2,000
Swadharma Topics
1M+
Pages of Śāstra Text
1,000
Hours of Lecture Video
The Knowledge Base

Four strata of the corpus

Every text enters the model through a graded trust system. The higher the stratum, the greater its weight in training — and every sentence remains traceable to the printed page it came from.

Gold

Editorially verified

The 72 volumes of the Veda Vijñāna Viṣṭaram Series — 15,000 pages of curated exposition spanning Veda, Vedāṅga, Itihāsa, Purāṇa, Āgama, Jyotiṣa, Sacred Geography, and the musical traditions of Bhārata. The supervised training spine of the model. Every volume, with its production status →

Silver

Curated & structured

The 2,000-topic Swadharma Master List; the SGSP Kalpa Sūtra Unicode collection — eleven Śrauta Sūtras, the Gṛhya register, commentated Dharma and Śulba texts — and curated assessment banks from SVAMI.

Bronze

Canonical, uncurated

Hundreds of scanned editions passing through the seven-stage industrial OCR pipeline: critical editions, traditional bhāṣyas of Sāyaṇa and the commentators, all eighteen Mahāpurāṇas, the Stotra literature, and the complete Kalpa Sūtra library in four classes.

Reference

Contextual

1,000 hours of transcribed lectures by traditional scholars, aligned with the text corpus — carrying the oral tradition's explanatory voice into the model.

The four strata above describe the corpus in principle. For what is actually in it — every text, its source edition, its trust tier and its verse count, regenerated whenever a module completes — see index.sarvaveda.info, the Master Index.

Ingested to date

43 texts, 374,439 addressable verses

All eighteen Mahāpurāṇas are in. Every figure below is counted from the corpus database, not entered by hand, and is regenerated whenever a module completes — Vedāṅga and Itihāsa and Purāṇa so far, 2 of the 11 departments.

A verse count marked * includes verses carrying a known doubt — 23 of the 43 texts. The doubt is named on the row: a generated script that will not convert back, a character no romanised Sanskrit should hold, a verse numbered by position because the source printed no number, or a locus the source repeats. The verses are all present and addressable; the asterisk says only that a reader should not treat the figure as settled.

Text and sourceCitationTierVerses in the text
Vedāṅga18 texts · 37,253 verses
Kalpa · Śrauta Sūtra11 · 30,326
Baudhāyana Śrauta SūtraBaudhSSsilver12,407SGSP Kalpa Sūtra Unicode collection
Āpastamba Śrauta SūtraApSSsilver8,331SGSP Kalpa Sūtra Unicode collection
Śāṅkhāyana Śrauta SūtraSankhSSsilver6,368SGSP Kalpa Sūtra Unicode collection
Bhāradvāja Śrauta SūtraBharSSsilver919SGSP Kalpa Sūtra Unicode collection
Mānava Śrauta SūtraManSSsilver802SGSP Kalpa Sūtra Unicode collection
Vārāha Śrauta SūtraVarSSsilver557SGSP Kalpa Sūtra Unicode collection
Maśaka Kalpa SūtraMasSSsilver310SGSP Kalpa Sūtra Unicode collection
Kātyāyana Śrauta SūtraKatySSsilver256SGSP Kalpa Sūtra Unicode collection
Āśvalāyana Śrauta SūtraAsvSSsilver212SGSP Kalpa Sūtra Unicode collection
Vaitāna Śrauta SūtraVaitSSsilver136SGSP Kalpa Sūtra Unicode collection
Jaiminīya Śrauta SūtraJaimSSsilver28SGSP Kalpa Sūtra Unicode collection
Kalpa · Gṛhya Sūtra7 · 6,927
Baudhāyana Gṛhya SūtraBaudhGSsilver3,303SGSP Kalpa Sūtra Unicode collection
Hiraṇyakeśin Gṛhya SūtraHirGSsilver1,193SGSP Kalpa Sūtra Unicode collection
Pāraskara Gṛhya SūtraParGSsilver813SGSP Kalpa Sūtra Unicode collection
Āśvalāyana Gṛhya SūtraAsvGSsilver762SGSP Kalpa Sūtra Unicode collection
Vaikhānasa Gṛhya SūtraVaikhGSsilver426SGSP Kalpa Sūtra Unicode collection
Āpastamba Gṛhya SūtraApGSsilver409SGSP Kalpa Sūtra Unicode collection
Khādira Gṛhya SūtraKhadGSsilver21SGSP Kalpa Sūtra Unicode collection
Itihāsa and Purāṇa25 texts · 337,186 verses
Itihāsa2 · 92,577
MahābhārataMBhsilver73,816BORI critical edition (Bhandarkar Oriental Research Institute, Pune) — John Smith's revision of M. Tokunaga
Vālmīki RāmāyaṇaRambronze18,761*Baroda critical edition text (M. Tokunaga, revised by John Smith), via GRETIL · * 32 whose generated script does not convert back · Kāvya by received classification; filed here with the Itihāsa on VKG's ruling
Mahāpurāṇa18 · 225,978
Padma Purāṇaपद्मपुराणम्PadPreference49,688*Saṃskṛta Wikisource dump, mainspace · * 269 whose generated script does not convert back; 666 numbered by position because the source printed none; 301 whose locus repeats in the source, filed as ~2
Brahmavaivarta Purāṇaब्रह्मवैवर्तपुराणम्BvPreference20,757*Saṃskṛta Wikisource dump, mainspace · * 982 whose generated script does not convert back; 21 numbered by position because the source printed none; 640 whose locus repeats in the source, filed as ~2
Bhaviṣya Purāṇaभविष्यपुराणम्BhavPreference18,629*Saṃskṛta Wikisource dump, mainspace · * 1,278 whose generated script does not convert back; 276 numbered by position because the source printed none; 208 whose locus repeats in the source, filed as ~2
Nārada PurāṇaNarPbronze15,586*Sansknet project · * 40 whose generated script does not convert back; 5 carrying characters no romanised Sanskrit should hold; 20 whose locus repeats in the source, filed as ~2 · holds Parts 1 and 2
Bhāgavata PurāṇaBhPbronze14,062*GRETIL per-skandha files · * 33 whose generated script does not convert back; 9 whose locus repeats in the source, filed as ~2 · holds Skandhas 1-12
Brahmāṇḍa PurāṇaBndPbronze13,743*Bombay: Veṅkaṭeśvara Steam Press · * 4 whose generated script does not convert back; 3 carrying characters no romanised Sanskrit should hold; 36 whose locus repeats in the source, filed as ~2 · holds 3 parts; verses 2,70.3-49 not available at present
Brahma PurāṇaBrPbronze13,579*Tübingen Purāṇa Project (Schreiner / Söhnen-Thieme) · * 19 whose generated script does not convert back; 2 carrying characters no romanised Sanskrit should hold; 1 whose locus repeats in the source, filed as ~2 · holds Adhyāyas 1-246
Garuḍa PurāṇaGarPbronze11,963*Bombay: Veṅkaṭeśvara Steam Press · * 3 whose generated script does not convert back; 1 carrying characters no romanised Sanskrit should hold; 10 whose locus repeats in the source, filed as ~2 · holds Parts 1-3
Agni PurāṇaAPbronze11,387*Bibliotheca Indica 65,1-3 · * 40 whose generated script does not convert back; 14 carrying characters no romanised Sanskrit should hold; 64 whose locus repeats in the source, filed as ~2
Varāha Purāṇaवराहपुराणम्VarPreference10,428*Saṃskṛta Wikisource dump, mainspace · * 225 whose generated script does not convert back; 5 numbered by position because the source printed none; 385 whose locus repeats in the source, filed as ~2
Liṅga PurāṇaLiPbronze9,135*Bombay: Veṅkaṭeśvara Steam Press 1906 · * 5 whose generated script does not convert back; 2 carrying characters no romanised Sanskrit should hold · holds Part 1 adhy. 1-108; part 2 adhy. 1-55
Matsya PurāṇaMatsPbronze8,498*Calcutta: Caukhamba Vidyabhavan 1954 · * 3 whose generated script does not convert back; 2 carrying characters no romanised Sanskrit should hold · holds Adhyāyas 1-176
Kūrma PurāṇaKūrmPbronze5,823*Sansknet project · * 1 whose locus repeats in the source, filed as ~2 · holds Parts 1 and 2
Vāmana PurāṇaVamPbronze5,683*A.S. Gupta (ed.), Varanasi: All India Kashiraj Trust 1967 · * 6 whose generated script does not convert back; 2 carrying characters no romanised Sanskrit should hold · holds Adhyāyas 1-69, plus the Saromāhātmya inserted after adhy. 23
Śiva PurāṇaŚivPbronze5,667*Bombay: Veṅkaṭeśvara Steam Press c. 1920 (input Jun Takashima) · * 15 carrying characters no romanised Sanskrit should hold · holds Book 1 Vidyeśvara-saṃhitā; book 7 Vāyavīya-saṃhitā parts 1-2
Viṣṇu PurāṇaViPbronze5,178*Critical edition, M.M. Pathak, Vadodara: Oriental Institute 1997-1999 · * 5 whose generated script does not convert back; 2 carrying characters no romanised Sanskrit should hold; 38 whose locus repeats in the source, filed as ~2
Mārkaṇḍeya PurāṇaMarkPbronze4,525*Sansknet project · * 7 carrying characters no romanised Sanskrit should hold · holds Adhyāyas 1-93 (94ff. not available at present)
Skanda PurāṇaSkPbronze1,647*Critical edition, Adriaensen / Bakker / Isaacson, Groningen: Egbert Forsten 1998- · * 1 whose generated script does not convert back · holds Adhyāyas 1-31.14 (to be continued)
Upapurāṇa3 · 4,134
Narasiṃha PurāṇaNsPbronze3,321*Siddheswar Jena (ed.), Delhi: Nag 1987 · * 8 whose generated script does not convert back; 78 carrying characters no romanised Sanskrit should hold
DevīgītāDgbronze510input Ursula Honegger · holds Devībhāgavata Purāṇa 7,31-40 only
Viṣṇudharmottara PurāṇaVdhabronze303*Bombay 1912 · * 2 whose generated script does not convert back; 2 carrying characters no romanised Sanskrit should hold; 2 whose locus repeats in the source, filed as ~2 · holds Adhyāyas 3,343-353 and 2,127 only
Khaṇḍa only2 · 14,497
Vāyu Purāṇa, RevākhaṇḍaRKVbronze7,776*Kṣemrāj Śrīkṛṣṇadās (ed.), Bombay 1910 · * 4 whose generated script does not convert back · holds Revākhaṇḍa — the only Vāyu material GRETIL holds
Skanda Purāṇa, RevākhaṇḍaRKSbronze6,721*Oṅkārānanda Giri (ed.), Hośaṅgābāda 1994 · * 1 whose generated script does not convert back · holds Revākhaṇḍa, complete

Last module loaded 2026-08-22. Every verse is addressable by its own citation — MBh.1.1.1, PadP.6.253.1 — and verifies byte-identical against the edition it came from. Module status, declared-against-ingested counts and every sub-project surface are at index.sarvaveda.info.

The Model Suite

Nine specialised engines, one foundation

A 70-billion-parameter foundation model, fine-tuned on the corpus and extended by nine śāstra-specific engines — each one auditable, each one citing its sources.

Śrauta Prayoga Generation

Complete rite sequences — ṛtvij assignments, mantra placement with svara, havirdravya, kāla-nirnaya — every step carrying a traceable sūtra citation.

Sāmaveda Stotra Generation

Stotriyā assembly, viṣṭuti patterns and stobha insertion for any Soma rite, in traditional numeric-svara and sargam notation.

Vikṛti Pāṭha Generation

All eight Vikṛti forms — Jaṭā to Ghana — for any Saṃhitā segment, with rule-verified sandhi and svara transformation at every junctura.

Tarka Sentence Structure

Pañcāvayava syllogisms, pūrvapakṣa–siddhānta frames, and Navya-Nyāya technical analysis of any śāstric argument.

Varṇa Krama Generation

Akṣara-by-akṣara decomposition with svara and mātrā, built directly on Śikṣā and Prātiśākhya rules — pāṭhaśālā-grade output.

Sāma Sound Analysis

Audio in, notation out: pitch and svara extraction from recorded Sāmagāna, automatic notation, and the Ūha engine that re-frames melody onto new chandas.

Vyākaraṇa Vocabulary Engine

The full Pāṇinian word-space — over 100 million forms, each with its complete sūtra-by-sūtra derivation — compiled into an instant-lookup lexicon.

Nirukta Etymology

Traditional nirvacana for any word, with competing etymologies cited, alongside the formal Pāṇinian vyutpatti — two tracks, one panel.

Chatur-Darśana Engine

The same sentence interpreted in parallel through Tarka, Vyākaraṇa, Mīmāṃsā and Vedānta — the śāstrārtha assembly, computed.

The Signature Capability

One sentence. Four śāstras.

In the traditional assembly, a naiyāyika, a vaiyākaraṇa, a mīmāṃsaka and a vedāntin each examine the same vākya with different instruments. SarvaVeda computes all four readings in parallel. Select a lens:

स्वर्गकामो यजेत
svargakāmo yajeta — "One desiring heaven should sacrifice"

The Nyāya reading

The sentence encodes a means–end inference: yāga is established as the sādhana for the sādhya, svarga. Reconstructed as a formal pañcāvayava: pratijñā — the desirer of heaven should perform yāga; hetu — because yāga is the instrument of heaven; udāharaṇa — whatever is an instrument to a desired end is to be undertaken by one who desires that end.

The analysis then tests the hetu for hetvābhāsa and states the cognition in Navya-Nyāya terms: svarga-niṣṭha-sādhyatā-nirūpita-sādhanatā resides in yāga.

Illustrative output — the production engine cites its sources for every claim.

Digitisation at Scale

The seven-stage OCR pipeline

Hundreds of scanned books — critical editions, commentaries, Grantha and Telugu texts — become verified, svara-correct digital corpus at a rhythm of 15–25 books a month.

O1
Intake & triage
O2
Image cleanup
O3
Layout analysis
O4
Script-aware OCR
O5
Svara recovery
O6
Lexicon correction
O7
Scholar gate

Two details set this pipeline apart. Vedic accent marks — which ordinary OCR silently discards — are recovered graphically or transferred from verified sources. And every OCR'd word is checked against the Vyākaraṇa engine's 100-million-form lexicon: the most complete Sanskrit spell-checker ever constructed, born from the same Pāṇinian generator that powers the grammar display.

The Plan

Six phases, thirty-six months

Foundation & infrastructure

Phase 0 · Months 1–4

Two GPU servers in operation on SGSP-controlled hardware, with vector and graph databases deployed and the scholar review panel constituted.

Corpus assembly

Phase 1 · Months 3–12

The 72 volumes, the Kalpa Sūtra Unicode collection, the Purāṇa and Stotra corpora ingested; the industrial OCR programme launched; the Vyākaraṇa lexicon generated; the citation register of every text quoted in the Series compiled.

Foundation model fine-tuning

Phase 2 · Months 8–16

QLoRA fine-tuning of the 70B base model on the Gold and Silver corpora; retrieval-augmented generation live; scholar-graded evaluation and feedback training.

The nine-engine suite

Phase 3 · Months 14–26

Prayoga, Stotra, Vikṛti, Tarka, Varṇa Krama, Nirukta and the Chatur-Darśana engines trained; the Sāma sound-analysis model delivers automatic notation; the web application opens for internal scholarly use.

Institutional pilot

Phase 4 · Months 20–28

SVAMI faculty and invited scholars test against the 2,000-topic benchmark; Prayoga outputs validated against the 2026 Agniṣṭoma performance records; the Ūha engine faces its Sāmaveda examiners.

Public release

Phase 5 · Months 26–36

Tiered public access — general, scholar, institutional — with the diaspora network of paṭhaśālās and centres as first subscribers, and a continuous-learning cycle absorbing each new volume and lecture.

Read the full 27-page roadmap → See the Plan of Action →
Custodians

The institutional foundation

SarvaVeda is not a standalone software project. It is the digital arm of a working Chaturveda Vidyālaya that has taught all four Vedas in the traditional gurukula manner since 2015 — the schools, faculty, admissions and projects of Veda Vijñāna Viṣṭaram are set out in full here.

SGSP

Sanatana Guru Sampradaya Pratishthanam, Mysore — the umbrella trust, and custodian of the VVV Series the corpus draws on.

Veda Vijñāna Viṣṭaram

The teaching and publishing body whose 72-volume Series forms the model's training spine — twelve schools, nine faculty, all four Vedas with their śākhās and Ghana Pāṭha. The Series →  ·  The institution →

SVAMI

The academy whose faculty serve as the scholar review panel — validating corpus quality and grading every model output.

Dharma Poshanam Inc.

USA 501(c)(3) — the vehicle for tax-deductible international support and diaspora institutional partnerships.

Build the model that serves the tradition

सत्यं ज्ञानमनन्तं ब्रह्म

SarvaVeda is seeking founding supporters, a Technical Director (AI/ML), and institutional partners. US donations are tax-deductible through Dharma Poshanam Inc. (501(c)(3)). Scholars, engineers, and patrons who wish to be part of this work are invited to write to the Pratishthanam.

Write to the project