Duhāī: eight books, counted
What eight compilations show that one could not
Take a paperback of spells, a machine-read copy of it, and a search box. That is the whole apparatus, and on this material it is enough to reach a wrong answer twice before it reaches a right one.
Start with the word आण, āṇ. Modern accounts of this material call it the defining technical move of the entire genre, the thing supposed to make a Śābara mantra a Śābara mantra rather than some other kind of spell. Type it. The file comes back with one hit.
One.
A finding of that shape writes itself. If the term everybody calls definitional stands once in a book of that length, then the popular account has taken a marginal register and promoted it into a definition.
The ghost is in the search box.
आण is one spelling. It carries the retroflex ṇ, and it is the spelling a Sanskritist reaches for first. The ordinary Hindī and Avadhi spelling is आन, with the dental n, and it is the spelling a reader of Tulsīdās or of Sūr would expect to meet. The same file, asked for आन, returns the word fifty-one times, and the count table below prints the two spellings together. Nothing in the single result was a fact about the language. It was a fact about a keystroke, and a silence manufactured by a keystroke reads exactly like a fact about a genre.
The second wrong answer sits one step further on, and it fails at the same joint. It says that the corpus prefers appeal to oath, that the two are distinguishable registers, and that the lexicographers put duhāī on the oath side of the line.
Open Śyāmasundara Dāsa’s Hindī Śabdasāgara, the standard dictionary of the language, compiled for the Nāgarī Pracāriṇī Sabhā at Kashi. Under दुहाई, in the original edition, the third sense of the word is given as शपथ, कसम, सोगंद: oath, swearing, vow. Under आन, in the revised edition, the second sense is oath and the third sense is proclamation, and that third sense is glossed by a single word. The word is दुहाई.
The two terms define each other. In the standard Indian lexicon they are near-twins, and one of them is used to explain the other. There is no appeal-versus-oath opposition in the language, and nothing can be built on one.
And one book counted is a case study of a book. What converts it into evidence is a second compilation of comparable size, counted the same way, with the two profiles printed side by side.
There are now eight. Seven more books, one script, one day, and the answer is bigger than a single volume could ever have shown it to be.
Eight books, and what they are a sample of
On 18 August 2026 the identical count was run over eight independent printed Śābara compilations. Six publishing houses. Three named cities. Just over 2.5 million characters of archive.org OCR, 479,254 Devanāgarī tokens, every one of them read from the same kind of machine text this treatise has been using all along, with no re-scanning and no hand correction.
The imprints stand as they are printed.
दुर्लभ शाबर मंत्र, Durlabh Śābar Mantra, presented by Tāntrik Bahal, Rājā Pocket Books of Burāṛī, Delhi 110084, new edition 2016, 323,692 characters and 62,014 tokens. That is the standing witness, the book chapters 10 to 13 are built on.
शाबरमन्त्रसागर, the Śābaramantrasāgara, first part, compiled by Śrī S. N. Khaṇḍelvāl, Chaukhambā Surabhāratī Prakāśana of Gopal Mandir Lane, Vārāṇasī 221001, a numbered volume in a Sanskrit series, edition of 2016, priced at ₹625, its compiler’s preface dated to Rāmanavamī 2013. 1,155,980 characters and 220,472 tokens.
प्राचीन सिद्ध शाबर मन्त्र, edited by Pramod Kumār Śāstrī, Rupeś Ṭhākur Prasād Prakāśan of Kachauṛī Galī, Vārāṇasī, printed at the Bhārat Press, dated on the imprint page to 2012. 348,185 characters, 58,029 tokens.
महाशक्तिशाली सिद्ध शाबर मंत्र, presented by M. I. Rājasvī, Pavan Pocket Books of Naī Saṛak, Delhi 110006, rights assigned on the copyright page to Gold Books (India), priced at ₹50, and carrying no date anywhere on the imprint page. 233,283 characters, 43,117 tokens.
दुर्लभ शाबर मंत्रों का रहस्य, Manoj Publications, undated, and an imperfect copy: the archive’s own record says pages are missing, so its absolute counts are a floor and only its rates are used. 123,831 characters, 24,182 tokens.
And three volumes of शाबर मन्त्र-संग्रह, parts 1, 3 and 7 of a numbered series, founding editor “Kula-bhūṣaṇ” Paṇḍit Ramādatta Śukla, M.A., of Prayāg, editor Ṛtaśīl Śarmā, published by Kalyāṇ Mandir Prakāśan of Alopī-Devī Mārg, Prayāg 211006, fourth edition, Gupta Navarātra of Saṃvat 2069, which is 20 June 2012, at ₹30 and ₹50. Together 391,342 characters and 71,440 tokens. Nine further parts of the same series are on archive.org, identified and left uncounted, because twelve volumes from one Allahabad house would have weighted the comparison toward one publisher.
Whether that Prayāg house has any relation to the Gita Press Kalyāṇ magazine of Gorakhpur is unverified. The names resemble each other. No connection is asserted here.
Now the question that has to be answered before any of the figures mean anything. A sample of what?
The population is still not bounded, and that has to be said before anything else. Nobody has produced a bibliography of the Hindī popular Śābara compilations, no union catalogue isolates them as a class, and this research still cannot say how many are in print. What has changed is that eight independent draws from that unbounded population now exist, made by six different houses, in three cities, across at least two quite different publishing economies. That is no longer a case study. It is a small sample with its selection method published, and its selection method was: everything the archive would give up under seven search strings, minus the manuscripts, minus the Sanskrit tantras, minus the duplicates.
One of the eight matters more than the others, and the reason is commercial rather than statistical. Chaukhambā Surabhāratī is a Sanskrit scholarly house at Vārāṇasī. Its Śābara volume carries a series number, an ISBN and a price of ₹625 against the ₹50 of the Delhi pocket books, and its compiler’s preface is sceptical in a way no bazaar paperback’s is: it observes that no Śābara material appears in the Siddhasiddhāntapaddhati and that the attribution of this corpus to Gorakhnāth has no evidence behind it. That preface is not cited here as an authority on anything, and this treatise takes no position from it. It matters for one reason only. If the distribution counted in a ₹50 Delhi paperback survives the jump to a ₹625 Banaras scholarly imprint whose own compiler doubts the standard sales pitch, the distribution is not an artefact of one publishing niche.
Everything below comes off one instrument, a string search over a machine-read text, and that instrument has a characteristic failure: a search run on a single orthographic form. It has produced miscounts on this material more than once, always from that same cause. What follows is published with the method attached, and with seven other instruments beside it.
Is it the scanner?
Before any of the comparison, one control the whole quantitative programme depends on.
The standing witness has been scanned twice. The same physical imprint, the same source PDF down to the byte size, uploaded to archive.org four and a half years apart and passed through two different releases of the same optical character recognition engine. The two files disagree about how much text is in the book, by eight per cent in the token count.
They do not disagree about the counts.
| what was counted | first scan | second scan |
|---|---|---|
| characters | 323,692 | 304,730 |
| Devanāgarī tokens | 62,014 | 56,821 |
| दुहाई duhāī | 71 | 72 |
| वाचा vācā | 64 | 65 |
| फुरो phuro | 49 | 46 |
| आज्ञा ājñā | 42 | 42 |
| आन ān, folded | 52 | 53 |
| पीर pīr | 18 | 19 |
| पैगम्बर paigambar | 2 | 2 |
| appeal frame, counted mechanically | 50 | 51 |
Two independent optical passes over one physical book move every headline figure by at most three, on a text whose extent they disagree about by eight per cent. Whatever else is wrong with these numbers, they are not artefacts of the scanner. That closes the scanner caveat, which no count run on a single scan can close.
The method, published
Reproducing anything below requires the following.
The text was taken from the full-text HTML of the archive.org item, with character entities unescaped and no tag-stripping pass applied afterwards. That order matters. Stripping pseudo-tags formed out of unescaped entities silently destroys about 85,000 characters of this file and returns a clean-looking and wrong total. The extraction that stands returns 323,692 characters, and that figure has been reproduced from two separate downloads of the same item. Two runs over one scan are still one scan, which is why the section above runs the figures against a second scan of the same book. The optical character recognition is poor: the file carries misrecognised Devanāgarī throughout, Latin garbage strings where the scanner met ornament, and at least six distinguishable misspellings of the single commonest proper name in the book.
The eight comparison texts were taken as the _djvu.txt layer archive.org publishes, which is the same optical layer the standing witness was read from, and the raw character figure quoted for each is the length of that file before anything is done to it. Every book was then normalised identically: Unicode NFC, zero-width non-joiner and zero-width joiner stripped, the nukta forms folded to their base letters and the free-standing combining nukta deleted, candrabindu folded to anusvāra. That third rule is not cosmetic on this material. The Perso-Arabic loan vocabulary in these books is printed with and without the nukta inconsistently inside a single volume, and the scanner adds a third layer of inconsistency on top. Nothing else was done. No case folding, which does not arise; no stemming; no normalising of long and short vowels, so that spellings which differ only in a vowel length remain separate strings and are folded by the variant list rather than by a rule.
Searches were run as plain string counts and, separately, as Devanāgarī token counts, with a token defined as a maximal run of characters in the Devanāgarī letter, matra and digit ranges, with daṇḍa, double daṇḍa and avagraha stripped from either end, so that space and punctuation break a token and a following vowel sign does not. The two rules do not agree with each other. वाचा returns 66 as a substring and 64 as a token; फुरो returns 54 and 49. Every term counted for this chapter was counted both ways, and the table below prints both figures on every row.
Those gaps are exactly where a table goes wrong. Put the substring figure for a term inside a table declared to hold token counts and the table is silently mixing two counting modes, and the mixing is invisible to anybody reading it. So every row of the table below is marked with its mode, because a frequency table with no counting mode against each row is not a published method. It is a number with a hidden rule behind it.
Orthographic variants were folded in for every counted term, without exception. Doing that moves several figures: वाचा gains बाचा, फुरो gains फूरो, आज्ञा gains आग्या, and आन gains आण. None of those is large. The point is the rule, which is that folding has to be done in both directions or neither, and cannot be applied to duhāī alone.
The fold lists used across all eleven books are these, matched as whole tokens. For the appeal word, दुहाई, दोहाई, दुहाइ, दुहाईं, दोहाइ and दुहायी. For ān, आन, आण, आंन and आनं. For vācā, वाचा and वांचा. For phuro, फुरो, फुरौ and the nukta spelling that the normalisation collapses into the first of them. For ājñā, आज्ञा, आग्या and आज्या. For pīr, पीर, पिर and four inflected forms. For paigambar, seven spellings across the nasal and the first vowel. For kasam, कसम with its oblique and plural forms and the nukta spelling. For saugandh, सौगंध, सौगन्ध and सौगंधि. A term counted on one spelling is a term counted on one publishing house’s compositor, and the section below on दोहाई shows what that costs.
The last of those foldings is the one most easily excepted. आन and आण are a dental and a retroflex spelling of the same word, and there is a narrative reason to print them as separate rows, because the retroflex is the spelling whose single occurrence produced the answer this chapter opened with. A narrative reason is not an exception to a rule, and a rule with an unstated exception in it is worse than no rule. The two are folded below into a single figure. The spelling that caused the trouble keeps its place in this chapter’s prose, where it belongs, and not a row of its own in a table that says variants have been folded.
Two of the figures below are structural rather than lexical, and how they are described matters more than what they are. One is a two-slot appeal frame, counted mechanically across every book by a rule that a replicator can re-derive and that is not restated here in a form that assembles it. The other is the number of distinct items standing in the authority slot of that frame. Those items were computed for every book, and not one of them is named anywhere in these chapters. The count is printed. The list is not, and it will not be, because printing the frame’s description and a run of its fillers on the same page is assembly by instalment, and a reader would be entitled to say so.
What is not published here is the hit lines. This treatise does not print working formulae, does not print their shape, and does not put a reader in a position to reconstruct one. Anyone auditing these numbers will ask for the hit lists, and the answer is that the lists of terms and of counts are printed in full below, the roster of names is printed as a roster, and the lines they were drawn from are not. That is the operative line, and it is not negotiable for a set of numbers.
What the Indian lexicon says
Take the three head-words to the dictionary that has the most Hindī in it.
Dāsa’s Hindī Śabdasāgara gives दुहाई three senses, with the etymology द्विधा + आह्वान. The first is proclamation, public crying, an announcement carried in all directions, illustrated from Sūr, from Jāyasī and from Tulsīdās. The second is the cry for help: calling out a name for rescue, and, in the compiler’s own framing, calling on someone of such power or eminence as can save when one is being tormented. The third is oath, swearing, vow, illustrated from Kabīr, from Sūr, from Tulsīdās and from Padmākar.
That second sense is the best gloss of the word this research has found anywhere, in any language. It names the social logic exactly: a person with a grievance, and a name invoked because its owner could intervene. It is more precise than anything in the English dictionaries, and it comes with dated literary attestation that the English dictionaries do not carry.
The third sense settles the question this treatise got wrong twice. Duhāī is itself an oath-word in Hindī, glossed by śapatha, kasam and saugandh.
Now आन, in the revised edition. The etymology given is आणि, limit or boundary. Sense 2 is oath, cited from the Rāmcaritmānas at Laṅkākāṇḍ 100. Sense 3 is proclamation of victory, glossed दुहाई, cited from Sūr. The remaining senses run through manner, haughtiness, deference, fear and vow.
The address for that entry is worth giving exactly, because it moves between printings. In the Nāgarī Pracāriṇī Sabhā revised-edition scan the आन entry stands between printed pages ४४४ and ४४५, and it was read there word for word. A citation a reader cannot follow is not a citation.
So āna is a Tulsī word and a Sūr word, and the standard dictionary of the language explains it with the other word. It is not a Rajasthani regionalism.
Mahendra Caturvedi’s A Practical Hindi-English Dictionary lists five senses of duhāī in this order: an outcry or entreaty for help, mercy or justice; plaint; oath; loud proclamation; and the process of or the wages paid for milking. Oath is third of five. John Platts, in his Dictionary of Urdu, Classical Hindi, and English, heads the entry with a cry for help, or mercy, or justice, complaint, exclamation, appeal, and then adds that the word is often used in obsecration, glossing rām-duhāʼī as a swearing by Rām.
Two readings of that evidence overshoot, in opposite directions. Quoting around “oath” hides a sense every one of the three dictionaries carries. Saying that all three class the word as an obsecration hides the other four. All three give appeal or proclamation as the primary sense and the oath sense as a further sense of the same word, and Platts’s etymology, a crying of hāy twice over, is an etymology of lament. No conclusion can be drawn from the lexica about a line between appeal and oath, because they do not draw one.
One note about which dictionary is doing what work. Dāsa leads on sense-range and on literary attestation, and is the better source on all three head-words tested. Turner’s Comparative Dictionary of the Indo-Aryan Languages and Platts are kept, in a narrow role, for comparative etymology: the Śabdasāgara’s derivations are traditional rather than comparative-historical, and द्विधा + आह्वान for duhāī is a folk derivation. Where this treatise needs Indo-Aryan descent or the Perso-Arabic stratum it goes to Turner and to Platts, and says so at the point of use. Where it needs the sense of a Hindī word it goes to Dāsa.
The count
One table, published once, for the standing witness. Every figure in the first column is a token count with orthographic variants folded, and the bracket is the fold. The second column is the substring count for the same folded set, given for every row and not only where it is convenient, because the gap between the two columns is the size of this instrument’s error bar. The third column notes what that gap is doing on each row.
| term | token | substring | note |
|---|---|---|---|
| दुहाई duhāī, with दोहाई and दुहाइ | 71 (59 + 11 + 1) | 71 | the two modes agree |
| वाचा vācā, with बाचा | 65 (64 + 1) | 68 | the two modes differ |
| आन ān, with आण | 52 (51 + 1) | 105 | the widest gap in the table |
| फुरो phuro, with फूरो | 51 (49 + 2) | 56 | the two modes differ |
| आज्ञा ājñā, with आग्या | 42 (40 + 2) | 44 | the two modes agree |
| शपथ śapatha | 1 | 1 | the two modes agree |
| कसम kasam | 0 | 1 | the substring is a different word |
| सौगंध saugandh | 0 | 0 | zero either way |
An independent re-run of that table, made on the same file with a script written from this chapter’s published rules, reproduces दुहाई, the आन fold and आज्ञा to the digit, and returns पैगम्बर, कीलित and उत्कीलन at two apiece. It reproduces वाचा and फुरो at the figures they carry before this chapter’s additional folds of बाचा and फूरो, which is the same answer arrived at by a different route. Six figures to the digit. That is what a published method is for.
The third column matters more than it looks. A table that mixes two counting modes without saying so does to its reader what a search on one spelling does to a genre: it produces a number whose rule is invisible, and an invisible rule cannot be argued with. The gaps involved are small, and being small is not the defence it looks like, because the direction of the error is unknowable from the printed figure alone.
Look at कसम. The token count is nought and the substring count is one, and that single substring hit is कसमसाने, a form of a different word altogether, meaning to stir or to writhe. That is precisely the false positive chapter 9 exposes for ताला, in the table whose entire argument is that this must not happen. A number produced by matching a string instead of a word.
Then the arithmetic of the columns. One bracket cannot carry two meanings, a token breakdown on one row and a substring total on the next, with no way for a reader to tell which is which. Every bracket above is a fold of token counts and sums to the token figure beside it, and the substring counts have a column of their own. The widest gap in the table is आन, and it is wide because those two letters open a great many longer Hindī words.
A machine count and a hand count are not measures of the same thing, and the honest figure is the one that says which it is. A search string measures a search string. A reader measures the book. Both have been run on the appeal word in the standing witness, and the answers differ, so they are set out here with the mode that produced each.
| measure, on दुहाई in the standing witness | mode | count |
|---|---|---|
| the word in one spelling, against one case-marker | machine, single string | 45 |
| the same test with the second spelling folded in | machine, folded strings | 48 |
| the appeal frame across the whole book | machine, frame rule | 50 |
| occurrences standing with a named authority | hand, one reader | ≈68 |
| occurrences of the word in the book | machine, folded tokens | 71 |
The distance between the machine figures and the hand figure is the part of the distribution a single string cannot reach, because this book uses more than one case-marker and no one pattern catches them all. Which marker stands where, and how the elements are ordered around it, is not reported here and will not be, in this chapter or any other. The hand check is one reader’s, on bad optical character recognition, and it should be redone by a second. The direction of the discrepancy is not in doubt. The smaller figures are facts about a search string, and the largest is a fact about the book.
Three figures are absent from the count table on purpose, and it is worth saying which and why.
The first is a matter of the operative line rather than of arithmetic. A two-word sequence recurring in this material returns thirty-three on the same instrument, and it has no row here. It is not a term. It is a construction, and this treatise does not print constructions from this material, or counts of them. A frequency table is an inventory, and an inventory of pieces is the one form in which this material could start to be assembled by somebody who wanted to assemble it. Chapter 11 takes the same item up as a philological problem, argues that it is most likely a corruption of an ordinary Sanskrit speech-frame with no connection to this genre at all, and builds nothing on it. That is the right place for it and this table is not.
The other two are matters of criterion. A figure of twenty-nine āna occurrences falling in unambiguous oath constructions is not printed, because it would rest on an unpublished criterion for what counts as an oath construction, because it is a positional claim of exactly the kind this treatise does not make about this material, and because no appeal-versus-oath opposition exists in the language for it to serve. Nor is a figure of roughly seven conditional-curse lines. There is no definition of “conditional curse” anywhere in this research, so such a number is not reproducible even in principle, and a number that cannot be reproduced is not evidence with a caveat attached. It is not evidence.
One thing the table deliberately does not do. It does not sort the terms into an oath group and an appeal group and invite a comparison of totals. Dāsa’s entries put duhāī and āna inside one another, and a table that separates them is drawing a line the dictionary refuses to draw. The columns above are token against substring, which is a fact about the instrument. They are not a fact about the language.
The same count, eight times
Absolute counts do not travel between books of different sizes, so everything from here is a rate: occurrences per 100,000 Devanāgarī tokens, with the absolute count in brackets behind it. Every figure is a token count. The bottom three rows are controls and are marked as controls.
| book | tokens | duhāī | appeal frame | vācā | phuro | ājñā |
|---|---|---|---|---|---|---|
| Rājā Pocket, Delhi 2016 | 62,014 | 114.5 (71) | 80.6 (50) | 103.2 (64) | 79.0 (49) | 67.7 (42) |
| Chaukhambā, Vārāṇasī 2016 | 220,472 | 299.4 (660) | 185.5 (409) | 291.6 (643) | 222.3 (490) | 49.9 (110) |
| Rupeś Ṭhākur, Vārāṇasī 2012 | 58,029 | 236.1 (137) | 160.3 (93) | 206.8 (120) | 167.2 (97) | 44.8 (26) |
| Pavan Pocket, Delhi n.d. | 43,117 | 136.8 (59) | 58.0 (25) | 41.7 (18) | 157.7 (68) | 16.2 (7) |
| Manoj, n.d., imperfect | 24,182 | 190.2 (46) | 124.1 (30) | 252.3 (61) | 272.9 (66) | 82.7 (20) |
| Kalyāṇ Mandir pt 1, Prayāg 2012 | 21,407 | 322.3 (69) | 79.4 (17) | 336.3 (72) | 210.2 (45) | 9.3 (2) |
| Kalyāṇ Mandir pt 3 | 16,362 | 299.5 (49) | 97.8 (16) | 213.9 (35) | 165.0 (27) | 18.3 (3) |
| Kalyāṇ Mandir pt 7 | 33,671 | 160.4 (54) | 65.3 (22) | 130.7 (44) | 68.3 (23) | 17.8 (6) |
| control: same compiler as row 1 | 22,104 | 280.5 (62) | 190.0 (42) | 131.2 (29) | 257.9 (57) | 58.8 (13) |
| control: Bṛhad Indrajāl 1964 | 30,053 | 10.0 (3) | 0 (0) | 10.0 (3) | 3.3 (1) | 10.0 (3) |
| control: Uḍḍīśatantra, Khemrāj | 13,183 | 0 (0) | 0 (0) | 0 (0) | 0 (0) | 0 (0) |
The appeal frame is the two-slot structure this chapter counts mechanically and does not describe. Its rule is a parser rule, it is reproducible from what is published here, and it is not written out as a sequence anywhere in these chapters.
Two reconciliations inside those figures are left open, and a reader weighing them should have both. The mechanical count of the appeal frame in the standing witness runs slightly above the lexical route, and the excess is most likely the second spelling folded into the frame test as well as into the lexical total. That is reported rather than resolved. And the hand-read figure stands well above what any machine rule returns. That is consistent with what is said above, that a hand reader reaches case-markers no single pattern catches. It is equally consistent with a hand reader counting generously. One reader on a poor scan cannot tell those apart, and the figure still needs a second reader.
Now read the duhāī column down.
The standing witness is the lowest of the eight on the appeal word. Not the middle. The lowest. Set it aside and every other book in the set carries that word more densely, the Banaras scholarly imprint at more than twice the Delhi paperback’s rate. Vācā moves the same way, and phuro, and the appeal frame.
A claim scoped to the standing witness alone would be too small.
That should be written plainly rather than celebrated. What the eight witnesses license is a scope of “the printed popular Śābara corpus, on eight independent witnesses from six houses,” in place of “this compilation,” and the volume these chapters happen to run on sits at the conservative end of the range. One witness carries a case study. Eight carry a claim, and the standing witness understates everything except one term.
The control that decides it
There is an obvious deflationary objection to everything above, and no amount of Śābara compilations answers it. It goes like this. Hindī bazaar occult printing is a genre with a house style. Of course these books are full of appeal words. They are cheap books about supernatural remedies, printed for the same readers, sold off the same racks, and any book of that kind would count the same way. The vocabulary is a property of the shelf.
So the identical script was run on two books from that shelf that are not Śābara compilations.
The first is Kautuk Ratna Bhaṇḍār, also titled Bṛhad Indrajāl, by Śyāmsundar Miśra, published by the N. S. Sharma Gaur Book Depot at Hāthras in 1964. Indrajāl is the neighbouring genre, the conjuring and marvels literature, and Hāthras is a bazaar press town of exactly the right kind. It stands in the table above, second from the bottom, and on the four headline terms it barely registers.
The second is the Uḍḍīśatantra with its Hindī ṭīkā by Paṇḍit Śyāmsundarlāl Tripāṭhī, from Khemrāj Śrīkṛṣṇadās at Mumbai, one of the imprints this treatise has been asking about since chapter 17. It is a tantra with a vernacular commentary, sold to the same customers. It stands on the bottom row. On the four headline terms it scores nought, nought, nought and nought.
Against 114 to 336 in the Śābara books.
That is a difference of one to two orders of magnitude, with a clean zero across the whole vocabulary in the second control. The appeal register is not a property of Hindī occult printing, of bazaar presses, of tantric subject matter, or of the scanning pipeline that produced every one of these files. It is a property of the Śābara compilations specifically, and it separates them from their nearest neighbours on the same shelf more cleanly than any other test in this treatise.
This is the strongest single result in the quantitative programme. Ten chapters have argued about what makes this material a category. Here is one thing that does, measured, with the negative controls printed beside the positives.
It is a lexical separation and nothing more. It does not show that the two genres do different things, that their readers wanted different things, or that the compilers thought of them as different kinds of book. It shows that four words which saturate one genre are almost absent from its neighbour.
Are these eight books, or one book eight times
A second objection, and it is the better one. Popular compilations copy each other. If these eight are one stock of material reprinted under eight covers, then eight witnesses are one witness with eight title pages, and the comparison above is arithmetic performed on a single text.
So the books were tested against each other. For every pair, the count of running seven-token strings they share, divided by the size of the smaller of the two. Optical noise depresses every one of those figures, so each is a floor rather than an estimate.
The standing witness and the Chaukhambā compilation share at most 4.6 per cent of their running seven-token strings. The comparison that carries this chapter is a comparison of two different books.
But the same test says something else, and it is more interesting than the reassurance. The Manoj book shares 28.6 per cent with the Chaukhambā compilation. The book by the standing witness’s own compiler shares 17.1 per cent with it. The Rupeś Ṭhākur book shares 9.4 per cent. The big Banaras compilation behaves like a reservoir that the smaller trade books draw on unevenly, and the corpus is a partially shared stock rather than either a set of independent compositions or a single text in eight jackets. Both halves of that finding are printed here because both are true and neither is convenient.
There is a third thing in the same table and it is the least comfortable of the three. The standing witness shares less with every other book than any other pair of books shares with each other. Of the eight, it is the odd one out. Chapters 10 to 13 run on the least representative compilation in the set, and the fact that its numbers converge with the others anyway is the reason the convergence means something.
The spelling that would have inverted the answer
The folding rule is the most-repeated methodological point in this chapter, and one book in the set shows exactly what it is worth.
Śābar Mantra-Saṅgrah part 1, the Kalyāṇ Mandir volume from Prayāg, writes the appeal word दोहाई five times out of six. Everywhere else in the set that is the minority spelling, and in the two largest books the ratio runs the other way.
| book | दुहाई | दोहाई | further variant | folded |
|---|---|---|---|---|
| Rājā Pocket, Delhi 2016 | 59 | 11 | 1 | 71 |
| Chaukhambā, Vārāṇasī 2016 | 544 | 116 | 0 | 660 |
| Kalyāṇ Mandir pt 1, Prayāg 2012 | 11 | 57 | 1 | 69 |
A researcher who searched the commoner spelling alone would have found eleven occurrences of the appeal word in the Prayāg volume, would have computed a rate of about fifty-one per 100,000, and would have written that the appeal vocabulary is thin in that series.
It is the densest book in the set.
Sixteen per cent recovery, and an inverted conclusion, from one orthographic choice. That is not methodological hygiene. On this corpus it is the difference between a right answer and its opposite.
The same test run on the ān fold produces the opposite error and is worth printing beside it. Across all eleven books the retroflex आण accounts for between nought and nine per cent of the fold, so searching that spelling alone recovers almost nothing, and searching the dental आन alone is very nearly right. Which is the exact inverse of the दुहाई case, and there is no way to know which situation you are in without running both.
What did not scale
Three findings run against the grain of everything above, and they are printed where they fall.
One belongs to chapter 11 and is stated here because the numbers live in this chapter. Ājñā, command or permission, is the headline term that goes down as the books get bigger. Read the ājñā column of the eight-book table from the top: the standing witness has the highest rate of the entire independent set, and the Banaras compilation, the Vārāṇasī book of 2012 and the Prayāg volumes fall away beneath it, the thinnest of them seven times thinner than the paperback. Ājñā is characteristic of Durlabh Śābar Mantra. It is not characteristic of the corpus. Chapter 11 treats it accordingly, as a feature of one book.
Another belongs to chapter 9, and it strengthens that chapter rather than weakening it. कीलित and उत्कीलन, the locking vocabulary, occur between nought and three times in every book counted, and the count does not move with the size of the book. The Delhi paperback and the Banaras compilation several times its length return the same figure, and the rest of the set varies inside that narrow band with no relation to extent at all. These are near-absent everywhere rather than rare-but-proportional, and that chapter’s finding that the locking discourse is marginal to this material is confirmed across eight independent witnesses and two genre controls. What may rest on the vocabulary is its presence, which is real and consistent. What no chapter may rest on is its frequency, because it does not have one.
And one is a clean negative. कसम and सौगंध, the Perso-Arabic oath vocabulary the कसमसाने false positive belongs to, are effectively absent from the whole corpus: nought tokens in nine of the eleven books, two and one in the other two. Whatever oath function these books perform, and the dictionaries say the appeal word performs one, it is not carried by the words Hindī normally uses for swearing. It is carried by duhāī and by the frame that stands with it. That is a negative result, it is worth as much as any of the positives, and it took eleven books to make it safe to state.
The names
Here is the roster. These are book-wide counts of the beings named in the standing witness, with obvious scanning variants folded in and the folding shown.
| named | count | folded across |
|---|---|---|
| Gorakhnāth | 379 | 366 standard, 13 in five distinguishable misscannings |
| Hanumān | 107 | |
| Īśvar | 80 | |
| Gaurā and Pārvatī | 58 | गौरा, गोरा, गौरी, गिरिजा, पार्वती |
| Bhairõ | 36 | |
| Mahādev | 33 | |
| pīr | 17 | of which more in chapter 12 |
| the Yoginīs | 16 | |
| the Nau Nāth and the Caurāsī Siddh | 15 | the two groups together |
| Kāmākṣā | 9 | |
| Rāmcandra | 8 | |
| Lonā Camārī | 6 | with Nonā Camārin |
| Kālkā | 6 | |
| Mohammadā | 4 | always as a vīr |
| Bhaĩsāsur | 3 | |
| Mansā | 2 | |
| Paigambar | 2 | |
| Viṣahārī | 1 | |
| Khwāja Khiḍr | 1 | |
| Autār | 1 |
Two things fall out of that list, and neither is visible without printing it.
This is a Gorakhnāth book. The lineage-founder outnumbers Śiva under the name Mahādev by more than eleven to one. Count only the occurrences captured by one search pattern and the two look level; that ranking is an artefact of the pattern.
And the goddess is not marginal in it. Gaurā and Pārvatī together stand at 58, ahead of Mahādev. Kāmākṣā, Kālkā, Mansā, Viṣahārī, the Yoginīs and Lonā Camārī are all present, and Lonā Camārī is the camār sorceress of North Indian folklore, a woman of the leatherworking caste standing in the position of invoked authority in a book that also names Śiva.
The roster is one reader’s hand-check of a bad scan. Sixteen name-families were folded; the folding rule was visual similarity plus contextual identity, and it was not applied blind. It needs a second reader and a better scan. What would close it: the same roster produced independently and the two lists compared.
One machine measure now stands beside the roster, and the two must not be confused. It counts the distinct items standing in the authority slot of the appeal frame.
| book | distinct items, by surface form | near-identical forms merged |
|---|---|---|
| Rājā Pocket, Delhi 2016 | 26 | 23 |
| Chaukhambā, Vārāṇasī 2016 | 119 | 95 |
Those are counts of forms. They are upper bounds by construction, and they are a different quantity from the roster above, which counts referents and merges every epithet and spelling of a being into a single line. A reader who sets the larger machine figure against the twenty names in the roster has set two incompatible measures against each other. Chapter 12 takes that figure up, and no name behind either count is printed in this treatise.
How these books are arranged
The eight tables of contents answer a question this treatise has carried unanswered since chapter 8, and the answer is not the one the classical literature would predict.
The six-act ṣaṭkarma grid does not organise these books. Not one of them.
Take the largest. The Chaukhambā compilation’s contents run continuously, and 1,451 entry titles were read off them and classified. Three hundred and twenty-five of those titles, 22.4 per cent, contain any six-act or ten-act term at all, and every one of those appears as an adjective inside the name of a single entry. None appears as a section division. Their internal distribution is not the classical six either.
| act named in an entry title | entries |
|---|---|
| immobilisation | 96 |
| subjugation | 80 |
| fascination | 54 |
| driving-away | 46 |
| killing | 18 |
| pacification | 16 |
| sowing-discord | 11 |
| attraction | 5 |
Two of the acts carry fifty-four per cent of all the six-act labels between them, and attraction, a canonical member of the set, appears five times across the whole of the largest Śābara compilation in print.
What organises the same 1,451 entries is the complaint.
| what the entry title names | entries |
|---|---|
| an affliction: an ailment, a pain, a trouble, a thing to be removed | 436 |
| money, trade, employment and fortune | 104 |
| possession, ghosts, witches and interference from outside | 99 |
| protection, armour and boundary-drawing | 86 |
| snakes, scorpions, dogs, cattle and vermin | 59 |
| theft and the recovery of stolen property | 16 |
Affliction alone is thirty per cent of the entries. The tail runs long and it runs domestic: toothache, jaundice, insomnia, hiccups, piles, ringworm, the bite of a rabid dog, a scorpion sting, a lost cow, a stalled business, the evil eye on a child.
The reader of the largest Śābara compilation in print is assumed to arrive with a problem. Not with a devotion, and not with a ritual programme.
Those are English category labels applied by one reader to entry titles read off a contents page, and the classification was by keyword. A second reader would move some entries between categories. The proportions are robust to that; the individual assignments are not.
Across the houses there are five architectures and the variation is itself the finding. The Chaukhambā book is a flat index of named purposes, clustering locally around a deity where the material happens to be deity-specific, in neighbourhoods within a list rather than chapters with headings. The Rupeś Ṭhākur book opens by declaring that human wants are met through daśa karma, ten acts, and lists and defines them one by one. That is the structural datum worth having: where these books reach for a classical grid at all, they reach for a ten-act one. The Manoj book divides therapy three ways, surgical, pharmaceutical and mantric, positions itself as the third for cases the first two fail on, and then runs by complaint under an explicitly medical frame. The Pavan Pocket book has no architecture beyond its promise, which is that the material needs no initiation and no guru, and it sits in a publisher’s back-list between a joke book and a gemstone guide.
And then there is the Prayāg series, which is organised by nobody’s system at all.
The Kalyāṇ Mandir volumes list their contents by contributor. Named correspondents, each with a home town: Jamnagar and Junagadh in Gujarat, Muzaffarnagar in Uttar Pradesh, Begusarai and Buxar in Bihar, Hamirpur in Himachal, Ratangarh in Churu district and Didwana in Rajasthan, Hadapsar in Pune, Sehore in Madhya Pradesh. Each contributor’s section is titled by what that person sent in. Three applications they had tried. Material obtained from sādhus and ojhās. A Śābara mantra in Punjabi. Rare applications. Only inside a contributor’s section is anything broken down by purpose or by deity. Part 1’s opening sections are editorial: a reading on Śābara power, an essay on the science of Śābara mantras by the founding editor, an essay on the role of the guru, a garland of the nine Nāths.
That is a periodical’s architecture inside book covers.
Twelve numbered parts, each a bound set of reader submissions, credited by name and by town across at least seven states, edited in Allahabad. It is a magazine. Any claim in this treatise about “the compiler’s choices” simply does not apply to three of the eight witnesses counted here, because there is no compiler there in the sense the claim needs. And any account of this material as a closed bazaar-press product manufactured in a shop has to reckon with a text visibly crowdsourced from practitioners across north and west India by an editor who printed their names and their addresses.
Contributor names are not printed here. Towns are, because the geography is the finding and the individuals are not.
The negative is as useful as the positive. Not one of the eight is organised alphabetically. Not one is organised by source-text. Not one is organised by occasion or by calendar. By purpose and by named affliction, overwhelmingly. By deity secondarily and locally. By contributor in one whole series. By a ten-act scheme where a classical grid is invoked at all. By the six-act ṣaṭkarma, never.
A parallel line of research in this project reached the same conclusion from the opposite direction, working from what the classical literature says the six acts are rather than from what the compilations do with them. Two methods, two starting points, one answer. The taxonomy argument is carried in a later chapter of this treatise, and what this section supplies to it is the contents-page evidence: eight tables of contents, five architectures, and the classical grid in none of them.
What the only scholarly survey actually says
Before the counting there is a prior question, and it is easy to answer by quoting one sentence.
Teun Goudriaan and Sanjukta Gupta’s Hindu Tantric and Śākta Literature, in Gonda’s History of Indian Literature, treats the Śābara Tantras in one short section. Quoted at one sentence, they appear to define the group by language register alone, with no speech-act in the definition at all.
Read the section instead of the sentence and that is false on the page.
Goudriaan and Gupta do say, in the sentence usually quoted, that the group is a small one and that its texts are mostly nothing but lists of spells, with Sanskrit and Hindī taking turns. That is a register description and it is real. The paragraph then keeps going, through the confusion of titles across the Śābara texts and an eighteenth-century argument by Kāśīnātha Bhaṭṭa for descent from the Kāpālikas, and it ends by stating what the group is principally about, which is getting command of the mantra that belongs to Śabara, whom they describe as a lesser deity with magical power, and the mantras of assorted other gods used for magical ends.
That is a subject-matter definition, sitting in the same short section as the register description, and it is easy to walk past.
The deity etymology is the part that matters. If Śabara names a minor deity of magical potency, then the whole “in barbarian tongue” reading this treatise has been building since chapter 2 has a live rival sitting in the one scholarly survey of the literature. It is not answered here. It is declared, and chapter 14 has to meet it.
What can be claimed from them is narrower and it is enough. Nothing in Goudriaan and Gupta, and nothing in any Tier 0 to Tier 2 source reached in this research, makes oath-binding definitional of this genre. Their two candidate definitions are one of register and one of subject matter. Neither is a speech-act.
What is actually being claimed
So here is the whole of it.
Across eight independently compiled printed Śābara books, from six houses in three cities, the most frequent recurrent lexical marker is a word for crying out in a name, and the rate table in this chapter gives its figure for every one of them. Beside it stand a word for speech and a verb form meaning let it come true, both of which rise with it. Two occult books from the same shelf that are not Śābara compilations score at or near zero on all three. The vocabulary of institutional qualification, the saṃskāra and the dīkṣā and the graded credential apparatus chapter 8 set out, is absent from all of them.
A fourth word, a noun meaning command or permission, does not belong with the others. It is dense in one book and thin in the rest, and chapter 11 says so.
That is a claim about a speech-act type, and it can be stated without any lexical opposition at all. What recurs across these compilations is a performative truth-assertion coupled to an appeal to a named authority. The speaker asserts that the word is true and stakes the assertion on somebody who is not the speaker. Whether the English for that is “appealing to” or “swearing by” is a question about English, and Dāsa’s dictionary declines to answer it because Hindī does not ask it.
Coercive address to divine and demonic powers is a real and well-documented category in charm literature elsewhere, and reading it here would not be an embarrassment. Jonathan Roper’s “Typologising English Charms,” in Charms and Charming in Europe (Palgrave Macmillan, London, 2004), supplies the comparative axis, sorting charms by whether they order, compare or ask. This chapter does not place the Śābara material on that axis, because placing it there would require reading constructions rather than counting words, and that is a thing this treatise has decided not to do in print.
What the counts refute is a definition. Not the presence of a register.
One more thing before this chapter hands over. Seven books and one script, against one book and a search box, did not simply confirm the smaller count. They raised its rate, put a second column on its table, took one of its four words away, and showed that the volume the whole quantitative programme runs on is the quietest in the set.
That is what a second witness is for.
The next chapter takes the other three words in the table, two of which the corpus confirms and one of which it takes away, and it takes them to the same dictionary, because the Śabdasāgara closes a philological question those words leave open.
They were not binding the gods, and they were not only pleading with them.
They were putting a name at stake, and hoping the name still carried.
Frequently asked
What does duhai mean?
Duhāī is glossed by Śyāmasundara Dāsa's Hindī Śabdasāgara in three senses: proclamation or public crying, the cry for help that calls on someone powerful enough to intervene, and oath, swearing or vow. The second sense names the social logic exactly, and the third is why the oath and appeal opposition does not survive.
Is a Shabara mantra defined by binding a deity with an oath?
The oath-binding definition is not supported by any of the eight compilations counted here, and no source at Tier 0 to Tier 2 reached in this research makes it definitional. Goudriaan and Gupta offer two candidate definitions, one of register and one of subject matter, and neither is a speech-act.
How many compilations have now been counted?
Eight independent printed compilations, from six houses in Delhi, Vārāṇasī and Prayāg, totalling 479,254 Devanāgarī tokens. Beside them stand three controls: a second scan of the first book, to measure how much of a count is optical artefact; a second book by the same compiler, to see whether one compiler reproduces one distribution; and two bazaar-press occult books that are not Śābara compilations, to see whether the vocabulary is generic to the genre of shop rather than to the genre of book.
Is the vocabulary specific to Shabara compilations?
Yes, and that is the sharpest result in the count. Bṛhad Indrajāl, Hāthras 1964, scores 10.0 per 100,000 tokens on the appeal word, 10.0 on vācā, 3.3 on phuro and nought on the appeal frame, and the Khemrāj Uḍḍīśatantra scores nought on all four, against 114 to 336 across the Śābara compilations. These are occult books from the same shelf and often the same shop. The gap is one to two orders of magnitude.
Why does the chapter insist on folding spelling variants?
Because on this corpus a single orthographic form can invert the answer. आण with retroflex ṇ returns one hit in the standing witness and आन with dental n returns fifty-one, folded to fifty-two. And Śābar Mantra-Saṅgrah part 1 spells the appeal word दोहाई fifty-seven times against दुहाई eleven, so a search on the commoner spelling alone would recover sixteen per cent of a book that is in fact the densest in the set.
Sources
- Durlabh Shabar Mantra, presented by Tantrik Bahal (Raja Pocket Books, Burari, Delhi, new edition 2016), archive.org item 20211212_20211212_0245, 323,692 characters; second scan of the same imprint at khta_durlabh-shabar-mantra-by-guru-gorakhnath-tantrik-bahal-hindi-tantra-illustr, 304,730 characters, used as an optical control.
- Shabaramantrasagara, purvabhaga, compiled by S. N. Khandelval (Chaukhamba Surabharati Prakashana, Varanasi, Surabharati Granthamala 568, edition 2016), ISBN 978-93-82443-95-7, contents indexed to p. 1475; archive.org item kjwd-shabar-mantra-sagar-s.-n.-khandelwal_202608, 1,155,980 characters.
- Pracina Siddha Shabara Mantra, ed. Pramod Kumar Shastri (Rupesh Thakur Prasad Prakashan, Varanasi, 2012), 348,185 characters; Mahashaktishali Siddha Shabar Mantra, presented by M. I. Rajasvi (Pavan Pocket Books, Delhi, undated), 233,283 characters; Durlabh Shabar Mantron ka Rahasya (Manoj Publications, undated, imperfect copy), 123,831 characters.
- Shabar Mantra-Sangrah, parts 1, 3 and 7, founding editor Ramadatta Shukla, editor Ritashil Sharma (Kalyan Mandir Prakashan, Prayag, 4th edition, Samvat 2069 Vi. = 20 June 2012); parts 2 and 4 to 12 located on archive.org and not counted.
- Genre controls: Kautuk Ratna Bhandar / Brihad Indrajal, Shyamsundar Mishra (N. S. Sharma Gaur Book Depot, Hathras, 1964), 30,053 tokens; Uddishatantra with Hindi tika, Pandit Shyamsundarlal Tripathi (Khemraj Shrikrishnadas Prakashan, Mumbai, 2017 printing, imperfect), 13,183 tokens. Dependent control: Siddha Shabar Mantra, Tantrik Bahal (Durga Pocket Books, Meerut, undated), 22,104 tokens.
- Shyamasundara Dasa, Hindi Shabdasagara (Nagari Pracharini Sabha, Kashi), duhai at original edition vol. 3 p. 1602; ana at revised edition Bhag 1 p. 444, the entry standing between printed pages 444 and 445 of the Nagari Pracharini Sabha revised-edition scan.
- Mahendra Chaturvedi, A Practical Hindi-English Dictionary, p. 325; John T. Platts, A Dictionary of Urdu, Classical Hindi, and English, p. 535.
- Teun Goudriaan and Sanjukta Gupta, Hindu Tantric and Sakta Literature (Harrassowitz, 1981), pp. 120-121; in copyright, paraphrased throughout.
- Jonathan Roper, "Typologising English Charms," in Charms and Charming in Europe (Palgrave Macmillan, London, 2004), pp. 128-144.
Read next
From the journal
Referenced in
Cited across 5 treatises · 1 reference entry.
The Prophets among the authorities
Treatise · chapterIslamic figures are named throughout the printed popular Śābara corpus. *Paigambar*, the word for prophet, stands twice in the Delhi pocket paperback…
LineagesThe philology of unintelligibility
Treatise · chapter*Anmil ākhar*, ill-matched letters, is best read as a verdict on how the syllables are dressed: badly assorted, wrong for the sound of sacred speech,…
LineagesThe Saora, who are still there
Treatise · chapterThe famous definition of *śābara* mantras, that they are addressed only to the deified ghosts of those who died violently, comes from a colonial…
LineagesThe word, reclaimed or not
Treatise · chapterThe word was never taken up by the people whose name it is. From the earliest lexical record to the marketing of 2026, *śābara* was repurposed by…
LineagesThe six acts
Treatise · chapterThe *ṣaṭkarma*, the six acts, is the grid the Sanskrit mantric compendia use to sort magical ritual, and it does not organise the Śābara corpus. In…
Reference · Sacred textShabara Tantra
A title belonging to no single book: an early Garuda scripture listed around 800 CE, and the late bilingual incantation corpus first printed at…
reference