SarvaVeda LLM ← Home
Vaidika Pada Kośa

Every word of the Veda, counted and addressed

The seven Padapāṭha texts of the four Vedas and two Brāhmaṇas have been read end to end, word by word. This is what they contain, what the reading found, and how each word is now citable.

Open the lexicon — search 1,286,475 words →

ॐ पदपाठः ॐ
463,281padas read
7Padapāṭha texts
4Vedas covered
13,099clusters
98.7%resolved by rule

Three treasuries of words

Rūḍhārtha · Vaidika · Yaugika

A word can be known in three different ways, and the tradition does not confuse them. The Pada Kośa keeps all three separately addressable.

A lexicographer’s entry is one kind of claim — citable to a printed page. A Pāṇinian derivation is another — citable to a sūtra. And a word actually attested in the Veda is a third: citable to its place in the recitation. Merging them would lose the one thing a Vedic scholar needs most, which is the ground of the meaning.

TreasuryDrawn fromAuthorityScale
Rūḍhārtha Pada Kośam62 Sanskrit lexiconsThe printed page312,682 words
Vaidika Pada Kośa7 Padapāṭha textsAttestation in the Veda463,281 padas
Yaugika Pada KośamThe AṣṭādhyāyīThe sūtra chain21 million forms

This page concerns the second. It is the newest of the three, and the only one whose contents are fixed by what the ṛṣis actually said, rather than by what a grammarian can derive or a lexicographer chose to record.

The census

A Padapāṭha is already word-separated — that is what it is for — so these are counts, not estimates. Counted by text rather than by printed volume, since volumes are a convenience of publishing.

TextPadasClustersCoverage
Ṛgveda Saṃhitā162,7161,992Complete
Kṛṣṇa Yajurveda — Taittirīya Saṃhitā110,0302,189Complete
Atharvaveda — Śaunaka Saṃhitā90,1706,050Complete
Taittirīya Brāhmaṇa43,361456Complete
Gopatha Brāhmaṇa32,571402Complete
Sāmaveda Saṃhitā24,4332,010Complete
Total463,28113,099

What the reading found

The Ṛgveda Padapāṭha is complete

Its two volumes are labelled 1–4 and 5–8, which invites the reading that maṇḍalas 9 and 10 are missing. They are not: the division is by aṣṭaka, and eight aṣṭakas are the whole Ṛgveda. The text says so in its own colophon —

इत्यष्टमाष्टकम्।। इत्यृग्वेद संहिता पदपाठः समाप्तः।।

“Thus the eighth aṣṭaka. Thus the Ṛgveda Saṃhitā Padapāṭha is complete.” The final verse carries the dual reference 8.8.49 = 10.190 — the last sūkta of the tenth maṇḍala.

Its verse count agrees with the tradition to 99.94%

Splitting the Ṛgveda on its verse terminator yields 10,558 mantras against the canonical 10,552 — arrived at independently of any external index, which is the strongest single confirmation that the text has been read correctly.

One text counts its words differently from the other six

The Gopatha Brāhmaṇa separates its padas with a comma where every other volume uses the daṇḍa — 32,127 commas against 1,198 daṇḍas. A reader that assumed the daṇḍa throughout would report the Gopatha as 434 words instead of 32,140, roughly one per cent of its true length, and would do so silently.

Words wrap across printed lines

In these typeset files a paragraph is a line, not a unit of text, and a single word is frequently cut across the break — सद्यता appearing as सद्य at the end of one line and ता at the start of the next. The Taittirīya Saṃhitā does this in 85% of its lines and the Gopatha in 99%. Counted naively, every such word is counted twice.

A gap worth recording. There is no Śukla Yajurveda Padapāṭha in the holding. Yajus coverage is Kṛṣṇa / Taittirīya only. The Kāṇva Saṃhitā is present, but as saṃhitā text rather than as a padapāṭha.

Two forms of every word

The word you recite and the word you analyse are not the same string, and the Kośa keeps both.

The Padapāṭha prints a word wrapped in its own analysis: a compound is given whole, then इति, then split into its members. That apparatus is essential for recitation and unusable for anything else. So each pada is held in more than one form:

FormWhat it isWhat it serves
OriginalVerbatim as printed, every accent intactPadapāṭha recitation
PrayogaThe whole word, peeled of its apparatusKrama and the eight Vikṛti pāṭhas, meaning, analysis
VeṣṭanaThe इति apparatus itselfLateral presentation beside the text

The peeling is exact for 98.7% of the corpus. Where it is not — a junction that needs sandhi, or a fusion whose original vowel cannot be recovered from the written form — the word is flagged for a scholar rather than guessed. Fifteen of the sixteen classical cases fall out of a single rule; the sixteenth requires vṛddhi, a change to the word itself rather than to its spelling, and is carried as a recorded exception.

The hierarchy in the source

The eight levels of the Vedic address — Veda, śākhā, grantha, aṣṭaka, adhyāya, sūkta, varga, vākya — are not imposed on these files. They are already marked within them, by dedicated symbols in the VijayaDV typeface, and can be read out mechanically.

LevelMarks the start ofAgreement with the canonical count
4Aṣṭaka / kāṇḍaExact — 8 aṣṭakas and 10 maṇḍalas share one symbol
5Adhyāya / prapāṭhaka96.9%
7Varga / pañcāśat / daśati97.9%
8Mantra vākya, and separately the Brāhmaṇa portions99.94% by verse terminator

That these marks are structural rather than decorative matters practically: a conversion that treated them as ornament would delete the hierarchy from all four Vedas while leaving the words themselves apparently intact.

Repetition

Punarukti — in the Ṛgveda Saṃhitā, गणान्तः — is the repetition of a phrase or mantra already given. It is marked in the text, and it is not a small feature: 1,439 marked spans carrying 7,976 padas, 4.90% of the Ṛgveda.

It must be tracked for two independent reasons. Krama and the Vikṛti pāṭhas cannot run across a repetition as though it were ordinary running text. And the Kośa must not record a repetition as a fresh attestation of a word, or the vocabulary inflates by every refrain in the Veda.

How a word is cited

Every addressable vākya already carries a fixed numeric address eight levels deep. The Padapāṭha adds a ninth — the word itself — without disturbing the existing eight:

01-01-01-01-001-01-01-001-0001  →  अग्निम्

Ṛgveda, first aṣṭaka, first adhyāya, first sūkta, first varga, first mantra, first word — the opening word of the Veda.

The ninth field counts within the cluster, so a word keeps its address even when the boundaries of a verse are redrawn. Measured across the corpus, the largest cluster holds 304 words against a field that accommodates 9,999 — ample room, established by measurement rather than by assumption.

Method and verification

Every figure on this page was measured against the sources. Where a count could be checked against the tradition’s own numbers, it was:

The text itself is held byte-exact. Vedic accents and the private-use characters of the VijayaDV typeface are never normalised, stripped, or silently converted at any stage, and transfers are checksummed to prove it.

Status. The reading and the word-level processing are complete for all seven texts. Full citation addressing is complete for the Ṛgveda and in progress for the remaining six. Nothing has been published to the corpus database, which remains under scholarly review.