Diary

Key developments, dated, with the workstream each belongs to.

A working record of the SarvaVeda project — what was built, when, and by which stream. Dates are taken from the repository's own history, so they are the dates the work actually landed, not the dates it was described.

22 July 2026  —  30 August 2026 · 21 entries

August 2026

  1. Four levels, three śāstras, and a fault that can be reported

    The app becomes a four-level workflow tree, and one sentence can be read by the śāstra whose question it is — Vyākaraṇa for the word, Mīmāṃsā for the sentence, Tarka for the warrant. A reader who finds a fault can report it without leaving the page, and it arrives as a numbered, labelled, assigned issue that cannot be closed without a named test and a run that passed. Elsewhere: 87 Phiṭsūtras where the tool had one, and 724 works on the register, each with three sentences that are evidence.

    ŚābdabodhaPada KośaCorpus
  2. A drill surface over 1,470 sentences

    The app is rebuilt over the corpus VKG supplied and audited: 49 constructs — 35 single words and 14 phrases — with 490 worked example sets, each climbing from the plain sentence to the fully abstracted pariṣkāra. Two navigation models are put up side by side, at / and at /alt, for scholars to choose between. see it

    Śābdabodha
  3. The Śābdabodha engine goes live

    Morphology over a 168,878-stem kośa, kāraka assignment by sūtra, compound classification, and the śābdabodha built three ways — Vyākaraṇa around the dhātvartha, Mīmāṃsā around the bhāvanā, Navya-Nyāya around the prathamānta. One relation is stated eight ways to show that the form is chosen, not forced. The three institutions are credited by role in every footer.

    ŚābdabodhaSite
  4. The Mīmāṃsā shelf captured

    Captured whole rather than a volume at a time, with every page cached so a re-run costs nothing, and the access token refreshed on the refusal rather than on a guessed timer.

    Corpus
  5. VKG's two axes, written as rules

    The categorising of Vedic sentences gets its definitions from the scholar and its enforcement from code: sentences are clubbed into vākyas on the understanding that a daṇḍa is not a full stop, and a broken sandhi prints as a dot rather than as a silent gap.

    MVVF
  6. The Master Index, and deploys that need no browser

    The whole Master Index is served, and shipping it becomes one command and then a git push from CI. What is live is provable against what was generated, by checksum, rather than by looking at it.

    InfrastructureSite
  7. A lexicon a researcher can actually use

    The Pada Kośa is built and published. Every word in the corpus is returned regardless of source; where a work is in copyright the card says to cite the printed edition, and only the attestation — which kośa records the word, and its printed page — is published. see it

    Pada Kośa
  8. 37,253 Kalpa Sūtras loaded, byte-identical

    Segmented at sūtra level, with the fonts Word records read alongside the text so the accents survive. The Itihāsa and Purāṇa haul is described in full. A position is taken that ārṣa literature is free unless a work forbids reuse in letter, rather than assuming restriction.

    Corpus
  9. The adhikaraṇas go up

    Published at /jaimini and /brahmasutra, browsable rather than only printable. see it

    SiteNyāyamālā
  10. 1,095 adhikaraṇas set as five volumes

    A4, 14pt, in VijayaDV. Every character is checked against the face before it is set, after a heading font turned out to carry 67 of the private-use glyphs where the text font has 1,143. Numerals are set in English figures, on VKG's ruling, so one font suffices alone.

    Nyāyamālā
  11. Both Nyāyamālās parsed into five-limb records

    The Jaiminīya and the Vaiyāsika are read into episode records carrying viṣaya, saṃśaya, pūrvapakṣa, siddhānta and saṅgati — the shape an adhikaraṇa is actually argued in.

    Nyāyamālā
  12. The plan reaches the eight working groups

    The register is published at /plan, the artefact identity ledger begins, and the plan is rewritten for each of the eight groups in language they can read. Deployment moves wholly to Cloudflare Pages. see it

    SiteCorpus
  13. Publishing becomes one command

    A Cloudflare Pages publisher for the site, and the Veda Vijñāna Series and the institution behind it brought onto sarvaveda.info. see it

    SiteInfrastructure
  14. Every unit on a dependency timeline

    A dependency chart that shows what blocks what, so the critical path is visible rather than argued about.

    Site
  15. Pada Kośa: three modules, and room to hold them

    The lexicon's architecture is settled as three modules, and the storage upgrade for the two compute nodes is specified against it.

    Pada KośaInfrastructure
  16. 31 volumes, 11,824 pages assembled

    SVU-027 (assemble the volumes already machine-readable) completes. The development plan is synced to issues so a unit's state is one place, and coordination tooling covers ownership, handover and absence.

    CorpusInfrastructure
  17. The Plan of Action is published

    Every document in the repository is numbered, the corpus registers are added, and the four-stage narrative goes up at /poa where it can be read rather than circulated. see it

    SiteCorpus

July 2026

  1. Phase 0 corpus audit, and the lexicon's two halves joined

    The data-holdings audit is written up as a consultation paper with its decisions and its risk register. Separately, the Pada Kośa's two compartments — the derived and the attested — are bridged, and the synset kośas parsed.

    CorpusPada Kośa
  2. The roadmap, deep

    A full /roadmap page, and a staged day-one bootstrap for the GPU machine so a new node can be brought up from a known state. see it

    SiteInfrastructure
  3. sarvaveda.info gets a face

    A landing page, the roadmap in full, and a pitch deck. The confidential finance page is deliberately kept out of the public repository from the first day rather than removed later. see it

    Site
  4. The repository is scaffolded

    Schema, development plan and the machine software stack are laid down together, with skeletons for ingestion and the API. The first ingestion run goes against real Veda Vijñāna Series data rather than a fixture, which is how the first defects were found.

    CorpusInfrastructure

The streams

Each entry belongs to one or more of these. They run in parallel and at different speeds; a quiet month in one is not a quiet month in the project.

CorpusAcquisition, OCR, segmentation and loading of the texts themselves.
Pada KośaThe lexicon — yaugika derivation and the rūḍha sources behind it.
NyāyamālāThe adhikaraṇa databases of the two Nyāyamālās, and their typesetting.
MVVFVeda Vākya Prakāratā Nirūpaṇam — the categorising of Vedic sentences.
ŚābdabodhaThe sentence-analysis engine and the app over it.
Sitesarvaveda.info and the published surfaces.
InfrastructureThe machines, the deployment path, and the tooling around both.
VKG Foundation — Content · Dharma Poshanam — Infrastructure · Veda Vijñāna Viṣṭaram — Evaluation