The seven Padapāṭha texts of the four Vedas and two Brāhmaṇas have been read end to end, word by word. This is what they contain, what the reading found, and how each word is now citable.
Open the lexicon — search 1,286,475 words →
Rūḍhārtha · Vaidika · Yaugika
A word can be known in three different ways, and the tradition does not confuse them. The Pada Kośa keeps all three separately addressable.
A lexicographer’s entry is one kind of claim — citable to a printed page. A Pāṇinian derivation is another — citable to a sūtra. And a word actually attested in the Veda is a third: citable to its place in the recitation. Merging them would lose the one thing a Vedic scholar needs most, which is the ground of the meaning.
| Treasury | Drawn from | Authority | Scale |
|---|---|---|---|
| Rūḍhārtha Pada Kośam | 62 Sanskrit lexicons | The printed page | 312,682 words |
| Vaidika Pada Kośa | 7 Padapāṭha texts | Attestation in the Veda | 463,281 padas |
| Yaugika Pada Kośam | The Aṣṭādhyāyī | The sūtra chain | 21 million forms |
This page concerns the second. It is the newest of the three, and the only one whose contents are fixed by what the ṛṣis actually said, rather than by what a grammarian can derive or a lexicographer chose to record.
A Padapāṭha is already word-separated — that is what it is for — so these are counts, not estimates. Counted by text rather than by printed volume, since volumes are a convenience of publishing.
| Text | Padas | Clusters | Coverage |
|---|---|---|---|
| Ṛgveda Saṃhitā | 162,716 | 1,992 | Complete |
| Kṛṣṇa Yajurveda — Taittirīya Saṃhitā | 110,030 | 2,189 | Complete |
| Atharvaveda — Śaunaka Saṃhitā | 90,170 | 6,050 | Complete |
| Taittirīya Brāhmaṇa | 43,361 | 456 | Complete |
| Gopatha Brāhmaṇa | 32,571 | 402 | Complete |
| Sāmaveda Saṃhitā | 24,433 | 2,010 | Complete |
| Total | 463,281 | 13,099 | — |
Its two volumes are labelled 1–4 and 5–8, which invites the reading that maṇḍalas 9 and 10 are missing. They are not: the division is by aṣṭaka, and eight aṣṭakas are the whole Ṛgveda. The text says so in its own colophon —
इत्यष्टमाष्टकम्।। इत्यृग्वेद संहिता पदपाठः समाप्तः।।
“Thus the eighth aṣṭaka. Thus the Ṛgveda Saṃhitā Padapāṭha is complete.” The final verse carries the dual reference 8.8.49 = 10.190 — the last sūkta of the tenth maṇḍala.
Splitting the Ṛgveda on its verse terminator yields 10,558 mantras against the canonical 10,552 — arrived at independently of any external index, which is the strongest single confirmation that the text has been read correctly.
The Gopatha Brāhmaṇa separates its padas with a comma where every other volume uses the daṇḍa — 32,127 commas against 1,198 daṇḍas. A reader that assumed the daṇḍa throughout would report the Gopatha as 434 words instead of 32,140, roughly one per cent of its true length, and would do so silently.
In these typeset files a paragraph is a line, not a unit of text, and a single word is frequently cut across the break — सद्यता appearing as सद्य at the end of one line and ता at the start of the next. The Taittirīya Saṃhitā does this in 85% of its lines and the Gopatha in 99%. Counted naively, every such word is counted twice.
A gap worth recording. There is no Śukla Yajurveda Padapāṭha in the holding. Yajus coverage is Kṛṣṇa / Taittirīya only. The Kāṇva Saṃhitā is present, but as saṃhitā text rather than as a padapāṭha.
The word you recite and the word you analyse are not the same string, and the Kośa keeps both.
The Padapāṭha prints a word wrapped in its own analysis: a compound is given whole, then इति, then split into its members. That apparatus is essential for recitation and unusable for anything else. So each pada is held in more than one form:
| Form | What it is | What it serves |
|---|---|---|
| Original | Verbatim as printed, every accent intact | Padapāṭha recitation |
| Prayoga | The whole word, peeled of its apparatus | Krama and the eight Vikṛti pāṭhas, meaning, analysis |
| Veṣṭana | The इति apparatus itself | Lateral presentation beside the text |
The peeling is exact for 98.7% of the corpus. Where it is not — a junction that needs sandhi, or a fusion whose original vowel cannot be recovered from the written form — the word is flagged for a scholar rather than guessed. Fifteen of the sixteen classical cases fall out of a single rule; the sixteenth requires vṛddhi, a change to the word itself rather than to its spelling, and is carried as a recorded exception.
The eight levels of the Vedic address — Veda, śākhā, grantha, aṣṭaka, adhyāya, sūkta, varga, vākya — are not imposed on these files. They are already marked within them, by dedicated symbols in the VijayaDV typeface, and can be read out mechanically.
| Level | Marks the start of | Agreement with the canonical count |
|---|---|---|
| 4 | Aṣṭaka / kāṇḍa | Exact — 8 aṣṭakas and 10 maṇḍalas share one symbol |
| 5 | Adhyāya / prapāṭhaka | 96.9% |
| 7 | Varga / pañcāśat / daśati | 97.9% |
| 8 | Mantra vākya, and separately the Brāhmaṇa portions | 99.94% by verse terminator |
That these marks are structural rather than decorative matters practically: a conversion that treated them as ornament would delete the hierarchy from all four Vedas while leaving the words themselves apparently intact.
Punarukti — in the Ṛgveda Saṃhitā, गणान्तः — is the repetition of a phrase or mantra already given. It is marked in the text, and it is not a small feature: 1,439 marked spans carrying 7,976 padas, 4.90% of the Ṛgveda.
It must be tracked for two independent reasons. Krama and the Vikṛti pāṭhas cannot run across a repetition as though it were ordinary running text. And the Kośa must not record a repetition as a fresh attestation of a word, or the vocabulary inflates by every refrain in the Veda.
Every addressable vākya already carries a fixed numeric address eight levels deep. The Padapāṭha adds a ninth — the word itself — without disturbing the existing eight:
01-01-01-01-001-01-01-001-0001 → अग्निम्
Ṛgveda, first aṣṭaka, first adhyāya, first sūkta, first varga, first mantra, first word — the opening word of the Veda.
The ninth field counts within the cluster, so a word keeps its address even when the boundaries of a verse are redrawn. Measured across the corpus, the largest cluster holds 304 words against a field that accommodates 9,999 — ample room, established by measurement rather than by assumption.
Every figure on this page was measured against the sources. Where a count could be checked against the tradition’s own numbers, it was:
The text itself is held byte-exact. Vedic accents and the private-use characters of the VijayaDV typeface are never normalised, stripped, or silently converted at any stage, and transfers are checksummed to prove it.
Status. The reading and the word-level processing are complete for all seven texts. Full citation addressing is complete for the Ṛgveda and in progress for the remaining six. Nothing has been published to the corpus database, which remains under scholarly review.