Instructions to use sapinsapin/halo-lid with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- fastText
How to use sapinsapin/halo-lid with fastText:
from huggingface_hub import hf_hub_download import fasttext model = fasttext.load_model(hf_hub_download("sapinsapin/halo-lid", "model.bin")) - Notebooks
- Google Colab
- Kaggle
halo-lid
A fastText language identifier for the ten languages of the Philippine
Language Dataset β bcl ceb eng fil hil ilo pag pam tsg war β plus other, trained on Indonesian and Malay
so the nearest non-Philippine neighbours are rejected rather than absorbed
into Cebuano.
Built for one job: gating a web scrape for Philippine-language text, where the alternatives either do not separate these ten languages well or are far too slow to run over millions of pages. 8.5 MB quantised, against GlotLID v3's 1.7 GB, and roughly 10x faster.
Results
Held-out set of 3,966 sentences from PLD read prompts β the only test set here whose labels were written by people rather than assigned by a model.
| model | accuracy | macro F1 | accuracy on <=5 words |
|---|---|---|---|
| halo-lid | 0.894 | 0.879 | 0.764 |
| GlotLID v3 | 0.705 | 0.752 | 0.428 |
Per language:
| lang | n | halo-lid | GlotLID v3 |
|---|---|---|---|
| bcl | 801 | 0.843 | 0.582 |
| ceb | 283 | 0.876 | 0.604 |
| eng | 417 | 0.969 | 0.842 |
| fil | 531 | 0.930 | 0.878 |
| hil | 238 | 0.824 | 0.676 |
| ilo | 634 | 0.923 | 0.781 |
| pag | 149 | 0.832 | 0.685 |
| pam | 582 | 0.950 | 0.692 |
| tsg | 107 | 0.822 | 0.439 |
| war | 224 | 0.804 | 0.598 |
Where it is weak. Waray (0.71) and Tausug (0.79) are the worst, and both regressed slightly from the previous round; their test sets are small (224 and 107 sentences), so treat the difference as indicative. Both are languages GlotLID is weak on too, which is part of the cause β see below.
How it was trained
scripts/train_lid.py in halohalo.
Training text: sentence-deduplicated PLD prompts (human labels), web text
from sapinsapin/halohalo, and pages accepted by our own scrape; negatives
are Indonesian and Malay from FineWeb-2.
The important detail is how web labels are handled. Web sentences inherit their page's language, which is wrong often enough to matter:
- Round 1 trained on page labels as-is and learned to call English
hil(English accuracy 0.63), because English sentences sit on Hiligaynon pages. - Round 2 relabelled confident English but still trusted the rest. It then
rejected 5,747 FineWeb-2 Filipino pages as Hiligaynon. The cause was
the training corpus: an audit of
sapinsapin/halo-hilfound GlotLID calls 44 % of a 2,000-sentence sample English, 21 % Filipino and only 12 % Hiligaynon β much of it Tagalog tabloid content. Round 2 had learned "Tagalog news is Hiligaynon". - Round 3 (this model) trains on a web sentence only when GlotLID agrees
with its page label (confident English is moved to
eng). That dropped 38khil-labelled sentences and fixed Filipino (0.759 -> 0.887). The known cost: it starves the languages GlotLID is weakest on, which is the likely reason Waray and Tausug slipped.
Usage
from halolib.lid import HaloLID # from the halohalo repo
HaloLID().predict("Maayong buntag sa inyong tanan") # -> ceb
Or with fastText directly:
import fasttext
from huggingface_hub import hf_hub_download
m = fasttext.load_model(hf_hub_download("sapinsapin/halo-lid", "model.ftz"))
m.predict("kumusta ka na kaibigan") # -> __label__fil
Input should be lowercased with digits and punctuation stripped; see
normalize() in halolib/lid.py. For documents, prefer
predict_document(), which votes over sentences weighted by length and also
returns an agreement score β useful for spotting code-switched pages.
Licence and limitations
cc-by-nc-4.0, because part of the training text comes from PLD, which is
CC-BY-NC and research-only. Treat this model as research-use.
othercovers only Indonesian and Malay. Spanish, Chavacano or other Philippine languages outside the ten will be forced into the nearest label, usually with low confidence.- Trained mostly on read prompts and web prose; performance on very short,
noisy, or heavily code-switched text (Taglish) is lower β the
<=5 wordscolumn is the honest indicator. filcovers Tagalog and Filipino together; PLD does not distinguish them.
- Downloads last month
- -