muril-hinglish-lid / README.md
pw-ai-research's picture
Initial release: MuRIL Hinglish LID tagger
78c7329 verified
|
Raw
History Blame Contribute Delete
10.7 kB
---
license: apache-2.0
base_model: google/muril-base-cased
language:
- hi
- en
library_name: transformers
pipeline_tag: token-classification
tags:
- token-classification
- language-identification
- code-switching
- hinglish
- code-mixed
- romanized-hindi
- text-to-speech
widget:
- text: "Mujhe calculus ka doubt samajh nahi aaya, please explain."
- text: "Main is question ka main point samajh nahi paya."
- text: "Is chapter mein bahut kam time hai, kam se kam revision kar lo."
model-index:
- name: muril-hinglish-lid
results:
- task:
type: token-classification
name: Language Identification (word-level)
dataset:
name: LinCE Hindi-English (dev, lang1+lang2 subset)
type: lince
metrics:
- type: accuracy
value: 0.9643
- type: f1
value: 0.9755
name: ENG F1
- type: f1
value: 0.9342
name: HIN F1
---
# muril-hinglish-lid
**Context-aware Hindi/English language identification for romanized Hinglish.**
Labels every word of code-mixed Latin-script text as `HIN` or `ENG`, using the
surrounding sentence rather than a dictionary.
## Example
Input:
```
Mujhe calculus ka doubt samajh nahi aaya, please explain.
```
Output:
| word | label |
|---|---|
| Mujhe | HIN |
| calculus | ENG |
| ka | HIN |
| doubt | ENG |
| samajh | HIN |
| nahi | HIN |
| aaya | HIN |
| please | ENG |
| explain | ENG |
## Why context matters
The same string can be two different languages in one sentence:
```
Main is question ka main point samajh nahi paya
HIN HIN ENG HIN ENG ENG HIN HIN HIN
└─ मैं └─ इस └─ English "main"
```
`Main` is मैं; four words later `main` is the English adjective. `is` is इस, not
the English copula. No lookup table can express this — it has one row per string.
That is not a contrived example. In our evaluation data the word `to` occurs 775
times, splitting **435 Hindi (तो) / 340 English**. A perfect lookup table gets at
most **56%** of those by always guessing the majority. This model gets **98.32%**.
## Usage
```python
import torch
from transformers import AutoTokenizer, AutoModelForTokenClassification
MODEL = "PhysicsWallahAI/muril-hinglish-lid"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForTokenClassification.from_pretrained(MODEL).eval()
def tag_words(words: list[str]) -> list[str]:
"""One label per word, read off the word's FIRST sub-token."""
enc = tok(words, is_split_into_words=True, truncation=True,
max_length=256, return_tensors="pt")
with torch.no_grad():
pred = model(**enc).logits[0].argmax(-1).tolist()
out, prev = [], None
for pos, wid in enumerate(enc.word_ids()):
if wid is not None and wid != prev:
out.append(model.config.id2label[pred[pos]])
prev = wid
return out
words = "Mujhe calculus ka doubt samajh nahi aaya please explain".split()
print(list(zip(words, tag_words(words))))
# [('Mujhe','HIN'), ('calculus','ENG'), ('ka','HIN'), ('doubt','ENG'),
# ('samajh','HIN'), ('nahi','HIN'), ('aaya','HIN'), ('please','ENG'),
# ('explain','ENG')]
```
**Use this helper rather than `pipeline(..., aggregation_strategy=...)`.** The
pipeline's aggregation is designed for NER: it merges *consecutive tokens sharing
a label* into one span, which is right for `New York City` → one `LOC` and wrong
here, where adjacent words routinely share a language. On the example above it
returns 7 spans instead of 9 words, fusing `samajh nahi aaya,` into a single
`HIN` blob. `aggregation_strategy="none"` is not the fix either — it returns
sub-word pieces (`'Mu'`, `'##jhe'`).
The helper matches the training-time contract exactly: words pre-split, one label
per word taken from its first sub-token. Continuation pieces were masked out of
the loss during training and carry no supervision, so reading them — or averaging
over them, which would let a long word's many pieces outvote a short word's one —
asks the model something it was never taught to answer.
Tag in windows of ~100 words, which is the window used in training.
## Model
- MuRIL backbone (`google/muril-base-cased`), pretrained on 17 Indian languages
**and their romanized forms**
- Token-classification head, 2 labels
- 237 M parameters · 12 layers · hidden 768 · WordPiece vocab 197,285
- Labels: `HIN` / `ENG`
- Context-aware prediction
## Performance
### Held-out gold set
874 answers / 108,432 tokens. The evaluation protocol is the part worth reading:
these labels were annotated **from the source text** by a separate model that
wrote none of the training corpus. Training labels were derived by aligning
Hinglish text to its Devanagari rewrite; the gold labels were not. So this is the
only measurement here that is independent of the pipeline that produced the
training data.
| metric | value |
|---|---|
| accuracy | **0.9920** |
| HIN F1 | 0.9927 |
| ENG F1 | 0.9912 |
| contextual homographs (8,444 occurrences) | 0.9861 |
| false-Hindi rate on English-only text | 0.02% (1 of 4,363 tokens) |
On its own auto-labelled test split the model reads 0.9934 — but that split is
the instrument that cannot see its own errors. The gold figure is the one to quote.
### Hard cases
| word | accuracy | note |
|---|---|---|
| `to` | 0.9832 | 775 occurrences, 435 तो / 340 English. Lookup ceiling: 56% |
| `use` | 0.7799 | उसे vs English "use". The weakest case, and unchanged across two independently retrained checkpoints — the signature of genuine ambiguity rather than a data defect |
| `beta` | — | बेटा (term of address) vs β, the physics symbol |
### Out of domain: LinCE Hindi-English
[LinCE](https://ritual.uh.edu/lince/) is human-annotated, public, and from a
different domain entirely — code-mixed **tweets**, not tutoring text.
| | tokens | accuracy | HIN F1 | ENG F1 |
|---|---|---|---|---|
| LinCE dev | 12,303 | **0.9643** | 0.9342 | 0.9755 |
> **Not comparable to the LinCE leaderboard.** LinCE has eight classes; this model
> has two. We map `lang1`→ENG, `lang2`→HIN and drop the rest — `other`
> (punctuation, handles, URLs, emoji; 2,231 tokens), `ne` (named entities; 875),
> and 37 tokens of `fw`/`mixed`/`unk`/`ambiguous`. Scoring only the two mappable
> classes is an **easier** task than LinCE's. Compare this number to our in-domain
> accuracy, and to nothing else.
### Latency
0.099 ms/word (108 words in 10.7 ms), NVIDIA A10, fp32, batch 1.
## Training
| | |
|---|---|
| data | Hinglish tutoring answers, labelled by aligning each answer to its Devanagari rewrite |
| train | 23,727 answers / 2,783,768 labelled tokens |
| dev | 1,402 answers / 162,047 labelled tokens |
| test | 2,771 answers / 322,569 labelled tokens |
| splits | **by answer** — the same answer never appears on two sides |
| hyperparameters | 3 epochs · batch 32 · lr 3e-5 · warmup ratio 0.1 · weight decay 0.01 · fp16 · max length 256 · 100-word windows |
| selection | best epoch by masked accuracy on dev; test read once, at the end |
Tokens the aligner could not place — math variables, symbols, ambiguous residue —
were emitted as `O` and **masked out of the loss** rather than guessed at, so the
model was never trained to invent a label for something the labelling process did
not actually know.
## Intended use
- Hinglish language identification
- Preprocessing for text-to-speech
- Transliteration pipelines
- Code-switched Hindi/English text
## Not intended for
- **General language identification across arbitrary languages.** Two labels
only. Other Indian languages in Latin script will be forced into `HIN` or `ENG`.
- **Devanagari transliteration itself.** This model decides *what* to
transliterate. It does not transliterate.
- **Determining whether a whole sentence is Hindi or English.** It is a
token-level model. Aggregating its labels to a sentence verdict is not what it
was built or evaluated for.
## Limitations
**SMS and chat shorthand.** The largest out-of-domain error class is abbreviated
social-media spelling, absent from tutoring text: `ur`, `u`, `h`, `k`, `b`, `r`,
`g`. On chat-register text expect worse than the LinCE number suggests — that
number already contains these errors, but diluted by well-formed tokens.
**Short fragments.** Accuracy tracks available context, not language. On short
homograph-dense English — *"Let me know so I can do it in the morning"* — error
rates rise sharply, because `let`, `me`, `know`, `so`, `can`, `do` are all Hindi
words in other contexts and six words condition almost nothing. On full-length
English text the false-Hindi rate is ~1%; on hand-picked short fragments it was
12.8%. Length is the variable.
**Named entities are out of scope.** The model emits only `HIN`/`ENG` and was
never trained on an NE class. Personal and place names receive *some* label, and
which one is not meaningful. Handle names separately.
**Domain.** Indian K-12 / exam-prep tutoring text: explanatory, second-person,
mathematics- and science-heavy, generally well-formed sentences. That is where
0.9920 holds.
## Data
The training data is proprietary tutoring text and is **not released**. No
training data is included in this repository — weights, config and tokenizer only.
The data reflects the register, subject matter and code-mixing conventions of one
setting: Hindi-English as used in Indian exam preparation. It is not a sample of
Hinglish in general, and the model's notion of "which words are Hindi" is that
community's.
Worth stating plainly, since this model exists to serve a TTS front-end: a `HIN`
label causes a word to be transliterated to Devanagari and pronounced with Hindi
phonology. A mislabelled word is therefore **mispronounced**, not dropped. The
failure is audible — the right direction for a failure to go, but it does mean
errors reach listeners directly.
## License
Apache-2.0, inherited from `google/muril-base-cased`. Training data is not
included and is not licensed for redistribution.
## Citation
```bibtex
@misc{muril-hinglish-lid,
title = {muril-hinglish-lid: context-aware Hindi/English language identification for romanized Hinglish},
author = {PhysicsWallah AI},
year = {2026},
url = {https://huggingface.co/PhysicsWallahAI/muril-hinglish-lid}
}
```
Base model:
```bibtex
@article{khanuja2021muril,
title = {MuRIL: Multilingual Representations for Indian Languages},
author = {Khanuja, Simran and Bansal, Diksha and Mehtani, Sarvesh and others},
journal = {arXiv preprint arXiv:2103.10730},
year = {2021}
}
```
Free AI Image Generator No sign-up. Instant results. Open Now