Skip to content

Improve Urdu language data and add a community training example - #14035

Open
mwzkhalil wants to merge 1 commit into
explosion:masterfrom
mwzkhalil:feat/urdu-language-support
Open

mwzkhalil wants to merge 1 commit into
explosion:masterfrom
mwzkhalil:feat/urdu-language-support

Conversation

@mwzkhalil

Copy link
Copy Markdown

Summary

  • Strengthen spacy.blank(ur) with copied tokenizer tables, Latin abbreviations, unspaced Urdu punctuation (including after harakat), and a like_num policy that treats U+060C as a list comma rather than a thousands separator.
  • Keep digit detection on Python str.isdigit for ASCII, Arabic-Indic, and Eastern Arabic-Indic digits; only واں/ویں forms with a known numeric stem count as numbers.
  • Add examples/urdu, a community training project with JSONL/CoNLL-U/BIO converters, official-split preservation, a CPU overfit diagnostic, and an optional XLM-R config. Licensed corpora and trained weights are not included.

This is language data and an example project, not an official Explosion Urdu pipeline.

Test plan

  • python -m pytest spacy/tests/lang/ur examples/urdu/tests -q
  • Confirm spacy.blank(ur) splits لفظ۔لفظ and ہےً۔کیا, keeps Dr. intact, and does not treat 1،2 or کتابواں as a single number
  • From examples/urdu, run python -m spacy project run overfit and treat the score as training-set learnability only
  • Confirm generated corpus/, training/, and *.spacy artefacts remain untracked
@mwzkhalil
mwzkhalil force-pushed the feat/urdu-language-support branch from 1cd9c18 to 6bb7861 Compare September 16, 2026 21:11

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

1 participant