Home / About
Mission

Spreading the pride of Tamil — without overreaching

Semmozhi (செம்மொழி, “classical language”) is for people who have never heard of Tamil as much as for those who grew up with it. The aim is simple: show how deep and how alive Tamil is, in a way a skeptical reader can check.

What we claim — and what we don’t

We do not say Tamil is “the first” or “the oldest” language in the world. That claim is rejected by independent fact-checkers and isn’t made by the scholars who know Tamil best. We stand behind the following, and each is being tied to a primary source:

ClaimWhere the evidence isStatus
Tamil is the only major Indian literary tradition independent of SanskritThe Tolkāppiyam grammar and Sangam poetry themselves; scholarship such as that of George L. HartAttach citations
2,000+ years of unbroken literary productionSangam anthologies → epics → devotional hymns → modern literature; see the readerAttach citations
60,000+ inscriptions — the densest epigraphic record in South AsiaEpigraphical Survey of India / Tamil Nadu archaeology recordsConfirm exact figure
A Roman-era maritime trade network with Tamil portsPeriplus of the Erythraean Sea (1st c. CE), the Muziris papyrus, excavations at ArikameduAttach citations
UNESCO’s World Heritage citation says “the Tamil civilisation”UNESCO — Great Living Chola Temples, Criterion (iii)Primary source linked

“Attach citations” means the claim comes from the project’s research plan and is well documented in the literature, but the exact footnotes haven’t been placed on this site yet. Do this before wide public launch.

Sources & licences

ContentSourceLicence
Chola & context articles (English)English Wikipedia; Wikivoyage (Thanjavur)CC BY-SA 4.0 — attribution shown on every article
Chola articles (Tamil)Tamil WikipediaCC BY-SA 4.0
Tirukkural, Tiruvācakam, Tirumantiram, Nālāyira Tivviya PirapantamProject Madurai e-textsFree for non-commercial distribution — keep attribution; contact Project Madurai for other uses
Sangam poetry, epics, grammar, NālaṭiyārTamil WikisourceCC BY-SA (individual works may carry their own notices — see the original page linked from each section)
Tēvāram (sample)thevaaram.orgCheck the site’s terms before reuse
ImagesWikimedia Commons (loaded from Commons, each linked to its author/licence page)Varies per image
Semmozhi Brahmi fontDerived from Noto Sans BrahmiSIL OFL 1.1
Brahmi conversion engineOriginal code, Unicode character namesYours to reuse

How this is built

The texts come from a crawl of open Tamil sources (about 28,700 records, 40 crawlers) that was classified by topic. Only a fraction is readable prose — the rest is catalogue data — so the site publishes what is actually complete and says plainly where it is partial. There is no server and no database: every page is static HTML plus JSON files, so it’s fast, cheap to host and easy to archive.

Known gaps

  • Partial texts. Aside from the Tirukkural, Tiruvācakam and parts of the Divya Prabandham, works are incomplete in the collected data. Each work is labelled Complete or Partial.
  • Pattuppāṭṭu (the Ten Idylls) are not in the corpus as primary text, so the site never claims “the full Sangam corpus”.
  • Chera, Pandya and Pallava kings lack deep profiles until better sources are gathered.
  • Script-evolution timeline is waiting for dedicated sources.
  • Wikipedia extracts are plain text — tables, infoboxes and citations are left out. Always follow the link to the original for those.