Spreading the pride of Tamil — without overreaching
Semmozhi (செம்மொழி, “classical language”) is for people who have never heard of Tamil as much as for those who grew up with it. The aim is simple: show how deep and how alive Tamil is, in a way a skeptical reader can check.
What we claim — and what we don’t
We do not say Tamil is “the first” or “the oldest” language in the world. That claim is rejected by independent fact-checkers and isn’t made by the scholars who know Tamil best. We stand behind the following, and each is being tied to a primary source:
| Claim | Where the evidence is | Status |
|---|---|---|
| Tamil is the only major Indian literary tradition independent of Sanskrit | The Tolkāppiyam grammar and Sangam poetry themselves; scholarship such as that of George L. Hart | Attach citations |
| 2,000+ years of unbroken literary production | Sangam anthologies → epics → devotional hymns → modern literature; see the reader | Attach citations |
| 60,000+ inscriptions — the densest epigraphic record in South Asia | Epigraphical Survey of India / Tamil Nadu archaeology records | Confirm exact figure |
| A Roman-era maritime trade network with Tamil ports | Periplus of the Erythraean Sea (1st c. CE), the Muziris papyrus, excavations at Arikamedu | Attach citations |
| UNESCO’s World Heritage citation says “the Tamil civilisation” | UNESCO — Great Living Chola Temples, Criterion (iii) | Primary source linked |
“Attach citations” means the claim comes from the project’s research plan and is well documented in the literature, but the exact footnotes haven’t been placed on this site yet. Do this before wide public launch.
Sources & licences
| Content | Source | Licence |
|---|---|---|
| Chola & context articles (English) | English Wikipedia; Wikivoyage (Thanjavur) | CC BY-SA 4.0 — attribution shown on every article |
| Chola articles (Tamil) | Tamil Wikipedia | CC BY-SA 4.0 |
| Tirukkural, Tiruvācakam, Tirumantiram, Nālāyira Tivviya Pirapantam | Project Madurai e-texts | Free for non-commercial distribution — keep attribution; contact Project Madurai for other uses |
| Sangam poetry, epics, grammar, Nālaṭiyār | Tamil Wikisource | CC BY-SA (individual works may carry their own notices — see the original page linked from each section) |
| Tēvāram (sample) | thevaaram.org | Check the site’s terms before reuse |
| Images | Wikimedia Commons (loaded from Commons, each linked to its author/licence page) | Varies per image |
| Semmozhi Brahmi font | Derived from Noto Sans Brahmi | SIL OFL 1.1 |
| Brahmi conversion engine | Original code, Unicode character names | Yours to reuse |
How this is built
The texts come from a crawl of open Tamil sources (about 28,700 records, 40 crawlers) that was classified by topic. Only a fraction is readable prose — the rest is catalogue data — so the site publishes what is actually complete and says plainly where it is partial. There is no server and no database: every page is static HTML plus JSON files, so it’s fast, cheap to host and easy to archive.
Known gaps
- Partial texts. Aside from the Tirukkural, Tiruvācakam and parts of the Divya Prabandham, works are incomplete in the collected data. Each work is labelled Complete or Partial.
- Pattuppāṭṭu (the Ten Idylls) are not in the corpus as primary text, so the site never claims “the full Sangam corpus”.
- Chera, Pandya and Pallava kings lack deep profiles until better sources are gathered.
- Script-evolution timeline is waiting for dedicated sources.
- Wikipedia extracts are plain text — tables, infoboxes and citations are left out. Always follow the link to the original for those.