Token Classification
Transformers
Safetensors
Hindi
English
bert
language-identification
code-switching
hinglish
code-mixed
romanized-hindi
text-to-speech
Eval Results (legacy)
Instructions to use PhysicsWallahAI/muril-hinglish-lid with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PhysicsWallahAI/muril-hinglish-lid with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="PhysicsWallahAI/muril-hinglish-lid")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("PhysicsWallahAI/muril-hinglish-lid") model = AutoModelForTokenClassification.from_pretrained("PhysicsWallahAI/muril-hinglish-lid", device_map="auto") - Notebooks
- Google Colab
- Kaggle
| license: apache-2.0 | |
| base_model: google/muril-base-cased | |
| language: | |
| - hi | |
| - en | |
| library_name: transformers | |
| pipeline_tag: token-classification | |
| tags: | |
| - token-classification | |
| - language-identification | |
| - code-switching | |
| - hinglish | |
| - code-mixed | |
| - romanized-hindi | |
| - text-to-speech | |
| widget: | |
| - text: "Mujhe calculus ka doubt samajh nahi aaya, please explain." | |
| - text: "Main is question ka main point samajh nahi paya." | |
| - text: "Is chapter mein bahut kam time hai, kam se kam revision kar lo." | |
| model-index: | |
| - name: muril-hinglish-lid | |
| results: | |
| - task: | |
| type: token-classification | |
| name: Language Identification (word-level) | |
| dataset: | |
| name: LinCE Hindi-English (dev, lang1+lang2 subset) | |
| type: lince | |
| metrics: | |
| - type: accuracy | |
| value: 0.9643 | |
| - type: f1 | |
| value: 0.9755 | |
| name: ENG F1 | |
| - type: f1 | |
| value: 0.9342 | |
| name: HIN F1 | |
| # muril-hinglish-lid | |
| **Context-aware Hindi/English language identification for romanized Hinglish.** | |
| Labels every word of code-mixed Latin-script text as `HIN` or `ENG`, using the | |
| surrounding sentence rather than a dictionary. | |
| ## Example | |
| Input: | |
| ``` | |
| Mujhe calculus ka doubt samajh nahi aaya, please explain. | |
| ``` | |
| Output: | |
| | word | label | | |
| |---|---| | |
| | Mujhe | HIN | | |
| | calculus | ENG | | |
| | ka | HIN | | |
| | doubt | ENG | | |
| | samajh | HIN | | |
| | nahi | HIN | | |
| | aaya | HIN | | |
| | please | ENG | | |
| | explain | ENG | | |
| ## Why context matters | |
| The same string can be two different languages in one sentence: | |
| ``` | |
| Main is question ka main point samajh nahi paya | |
| HIN HIN ENG HIN ENG ENG HIN HIN HIN | |
| └─ मैं └─ इस └─ English "main" | |
| ``` | |
| `Main` is मैं; four words later `main` is the English adjective. `is` is इस, not | |
| the English copula. No lookup table can express this — it has one row per string. | |
| That is not a contrived example. In our evaluation data the word `to` occurs 775 | |
| times, splitting **435 Hindi (तो) / 340 English**. A perfect lookup table gets at | |
| most **56%** of those by always guessing the majority. This model gets **98.32%**. | |
| ## Usage | |
| ```python | |
| import torch | |
| from transformers import AutoTokenizer, AutoModelForTokenClassification | |
| MODEL = "PhysicsWallahAI/muril-hinglish-lid" | |
| tok = AutoTokenizer.from_pretrained(MODEL) | |
| model = AutoModelForTokenClassification.from_pretrained(MODEL).eval() | |
| def tag_words(words: list[str]) -> list[str]: | |
| """One label per word, read off the word's FIRST sub-token.""" | |
| enc = tok(words, is_split_into_words=True, truncation=True, | |
| max_length=256, return_tensors="pt") | |
| with torch.no_grad(): | |
| pred = model(**enc).logits[0].argmax(-1).tolist() | |
| out, prev = [], None | |
| for pos, wid in enumerate(enc.word_ids()): | |
| if wid is not None and wid != prev: | |
| out.append(model.config.id2label[pred[pos]]) | |
| prev = wid | |
| return out | |
| words = "Mujhe calculus ka doubt samajh nahi aaya please explain".split() | |
| print(list(zip(words, tag_words(words)))) | |
| # [('Mujhe','HIN'), ('calculus','ENG'), ('ka','HIN'), ('doubt','ENG'), | |
| # ('samajh','HIN'), ('nahi','HIN'), ('aaya','HIN'), ('please','ENG'), | |
| # ('explain','ENG')] | |
| ``` | |
| **Use this helper rather than `pipeline(..., aggregation_strategy=...)`.** The | |
| pipeline's aggregation is designed for NER: it merges *consecutive tokens sharing | |
| a label* into one span, which is right for `New York City` → one `LOC` and wrong | |
| here, where adjacent words routinely share a language. On the example above it | |
| returns 7 spans instead of 9 words, fusing `samajh nahi aaya,` into a single | |
| `HIN` blob. `aggregation_strategy="none"` is not the fix either — it returns | |
| sub-word pieces (`'Mu'`, `'##jhe'`). | |
| The helper matches the training-time contract exactly: words pre-split, one label | |
| per word taken from its first sub-token. Continuation pieces were masked out of | |
| the loss during training and carry no supervision, so reading them — or averaging | |
| over them, which would let a long word's many pieces outvote a short word's one — | |
| asks the model something it was never taught to answer. | |
| Tag in windows of ~100 words, which is the window used in training. | |
| ## Model | |
| - MuRIL backbone (`google/muril-base-cased`), pretrained on 17 Indian languages | |
| **and their romanized forms** | |
| - Token-classification head, 2 labels | |
| - 237 M parameters · 12 layers · hidden 768 · WordPiece vocab 197,285 | |
| - Labels: `HIN` / `ENG` | |
| - Context-aware prediction | |
| ## Performance | |
| ### Held-out gold set | |
| 874 answers / 108,432 tokens. The evaluation protocol is the part worth reading: | |
| these labels were annotated **from the source text** by a separate model that | |
| wrote none of the training corpus. Training labels were derived by aligning | |
| Hinglish text to its Devanagari rewrite; the gold labels were not. So this is the | |
| only measurement here that is independent of the pipeline that produced the | |
| training data. | |
| | metric | value | | |
| |---|---| | |
| | accuracy | **0.9920** | | |
| | HIN F1 | 0.9927 | | |
| | ENG F1 | 0.9912 | | |
| | contextual homographs (8,444 occurrences) | 0.9861 | | |
| | false-Hindi rate on English-only text | 0.02% (1 of 4,363 tokens) | | |
| On its own auto-labelled test split the model reads 0.9934 — but that split is | |
| the instrument that cannot see its own errors. The gold figure is the one to quote. | |
| ### Hard cases | |
| | word | accuracy | note | | |
| |---|---|---| | |
| | `to` | 0.9832 | 775 occurrences, 435 तो / 340 English. Lookup ceiling: 56% | | |
| | `use` | 0.7799 | उसे vs English "use". The weakest case, and unchanged across two independently retrained checkpoints — the signature of genuine ambiguity rather than a data defect | | |
| | `beta` | — | बेटा (term of address) vs β, the physics symbol | | |
| ### Out of domain: LinCE Hindi-English | |
| [LinCE](https://ritual.uh.edu/lince/) is human-annotated, public, and from a | |
| different domain entirely — code-mixed **tweets**, not tutoring text. | |
| | | tokens | accuracy | HIN F1 | ENG F1 | | |
| |---|---|---|---|---| | |
| | LinCE dev | 12,303 | **0.9643** | 0.9342 | 0.9755 | | |
| > **Not comparable to the LinCE leaderboard.** LinCE has eight classes; this model | |
| > has two. We map `lang1`→ENG, `lang2`→HIN and drop the rest — `other` | |
| > (punctuation, handles, URLs, emoji; 2,231 tokens), `ne` (named entities; 875), | |
| > and 37 tokens of `fw`/`mixed`/`unk`/`ambiguous`. Scoring only the two mappable | |
| > classes is an **easier** task than LinCE's. Compare this number to our in-domain | |
| > accuracy, and to nothing else. | |
| ### Latency | |
| 0.099 ms/word (108 words in 10.7 ms), NVIDIA A10, fp32, batch 1. | |
| ## Training | |
| | | | | |
| |---|---| | |
| | data | Hinglish tutoring answers, labelled by aligning each answer to its Devanagari rewrite | | |
| | train | 23,727 answers / 2,783,768 labelled tokens | | |
| | dev | 1,402 answers / 162,047 labelled tokens | | |
| | test | 2,771 answers / 322,569 labelled tokens | | |
| | splits | **by answer** — the same answer never appears on two sides | | |
| | hyperparameters | 3 epochs · batch 32 · lr 3e-5 · warmup ratio 0.1 · weight decay 0.01 · fp16 · max length 256 · 100-word windows | | |
| | selection | best epoch by masked accuracy on dev; test read once, at the end | | |
| Tokens the aligner could not place — math variables, symbols, ambiguous residue — | |
| were emitted as `O` and **masked out of the loss** rather than guessed at, so the | |
| model was never trained to invent a label for something the labelling process did | |
| not actually know. | |
| ## Intended use | |
| - Hinglish language identification | |
| - Preprocessing for text-to-speech | |
| - Transliteration pipelines | |
| - Code-switched Hindi/English text | |
| ## Not intended for | |
| - **General language identification across arbitrary languages.** Two labels | |
| only. Other Indian languages in Latin script will be forced into `HIN` or `ENG`. | |
| - **Devanagari transliteration itself.** This model decides *what* to | |
| transliterate. It does not transliterate. | |
| - **Determining whether a whole sentence is Hindi or English.** It is a | |
| token-level model. Aggregating its labels to a sentence verdict is not what it | |
| was built or evaluated for. | |
| ## Limitations | |
| **SMS and chat shorthand.** The largest out-of-domain error class is abbreviated | |
| social-media spelling, absent from tutoring text: `ur`, `u`, `h`, `k`, `b`, `r`, | |
| `g`. On chat-register text expect worse than the LinCE number suggests — that | |
| number already contains these errors, but diluted by well-formed tokens. | |
| **Short fragments.** Accuracy tracks available context, not language. On short | |
| homograph-dense English — *"Let me know so I can do it in the morning"* — error | |
| rates rise sharply, because `let`, `me`, `know`, `so`, `can`, `do` are all Hindi | |
| words in other contexts and six words condition almost nothing. On full-length | |
| English text the false-Hindi rate is ~1%; on hand-picked short fragments it was | |
| 12.8%. Length is the variable. | |
| **Named entities are out of scope.** The model emits only `HIN`/`ENG` and was | |
| never trained on an NE class. Personal and place names receive *some* label, and | |
| which one is not meaningful. Handle names separately. | |
| **Domain.** Indian K-12 / exam-prep tutoring text: explanatory, second-person, | |
| mathematics- and science-heavy, generally well-formed sentences. That is where | |
| 0.9920 holds. | |
| ## Data | |
| The training data is proprietary tutoring text and is **not released**. No | |
| training data is included in this repository — weights, config and tokenizer only. | |
| The data reflects the register, subject matter and code-mixing conventions of one | |
| setting: Hindi-English as used in Indian exam preparation. It is not a sample of | |
| Hinglish in general, and the model's notion of "which words are Hindi" is that | |
| community's. | |
| Worth stating plainly, since this model exists to serve a TTS front-end: a `HIN` | |
| label causes a word to be transliterated to Devanagari and pronounced with Hindi | |
| phonology. A mislabelled word is therefore **mispronounced**, not dropped. The | |
| failure is audible — the right direction for a failure to go, but it does mean | |
| errors reach listeners directly. | |
| ## License | |
| Apache-2.0, inherited from `google/muril-base-cased`. Training data is not | |
| included and is not licensed for redistribution. | |
| ## Citation | |
| ```bibtex | |
| @misc{muril-hinglish-lid, | |
| title = {muril-hinglish-lid: context-aware Hindi/English language identification for romanized Hinglish}, | |
| author = {PhysicsWallah AI}, | |
| year = {2026}, | |
| url = {https://huggingface.co/PhysicsWallahAI/muril-hinglish-lid} | |
| } | |
| ``` | |
| Base model: | |
| ```bibtex | |
| @article{khanuja2021muril, | |
| title = {MuRIL: Multilingual Representations for Indian Languages}, | |
| author = {Khanuja, Simran and Bansal, Diksha and Mehtani, Sarvesh and others}, | |
| journal = {arXiv preprint arXiv:2103.10730}, | |
| year = {2021} | |
| } | |
| ``` | |