halo-lid

A fastText language identifier for the ten languages of the Philippine Language Dataset β€” bcl ceb eng fil hil ilo pag pam tsg war β€” plus other, trained on Indonesian and Malay so the nearest non-Philippine neighbours are rejected rather than absorbed into Cebuano.

Built for one job: gating a web scrape for Philippine-language text, where the alternatives either do not separate these ten languages well or are far too slow to run over millions of pages. 8.5 MB quantised, against GlotLID v3's 1.7 GB, and roughly 10x faster.

Results

Held-out set of 3,966 sentences from PLD read prompts β€” the only test set here whose labels were written by people rather than assigned by a model.

model accuracy macro F1 accuracy on <=5 words
halo-lid 0.894 0.879 0.764
GlotLID v3 0.705 0.752 0.428

Per language:

lang n halo-lid GlotLID v3
bcl 801 0.843 0.582
ceb 283 0.876 0.604
eng 417 0.969 0.842
fil 531 0.930 0.878
hil 238 0.824 0.676
ilo 634 0.923 0.781
pag 149 0.832 0.685
pam 582 0.950 0.692
tsg 107 0.822 0.439
war 224 0.804 0.598

Where it is weak. Waray (0.71) and Tausug (0.79) are the worst, and both regressed slightly from the previous round; their test sets are small (224 and 107 sentences), so treat the difference as indicative. Both are languages GlotLID is weak on too, which is part of the cause β€” see below.

How it was trained

scripts/train_lid.py in halohalo. Training text: sentence-deduplicated PLD prompts (human labels), web text from sapinsapin/halohalo, and pages accepted by our own scrape; negatives are Indonesian and Malay from FineWeb-2.

The important detail is how web labels are handled. Web sentences inherit their page's language, which is wrong often enough to matter:

  • Round 1 trained on page labels as-is and learned to call English hil (English accuracy 0.63), because English sentences sit on Hiligaynon pages.
  • Round 2 relabelled confident English but still trusted the rest. It then rejected 5,747 FineWeb-2 Filipino pages as Hiligaynon. The cause was the training corpus: an audit of sapinsapin/halo-hil found GlotLID calls 44 % of a 2,000-sentence sample English, 21 % Filipino and only 12 % Hiligaynon β€” much of it Tagalog tabloid content. Round 2 had learned "Tagalog news is Hiligaynon".
  • Round 3 (this model) trains on a web sentence only when GlotLID agrees with its page label (confident English is moved to eng). That dropped 38k hil-labelled sentences and fixed Filipino (0.759 -> 0.887). The known cost: it starves the languages GlotLID is weakest on, which is the likely reason Waray and Tausug slipped.

Usage

from halolib.lid import HaloLID          # from the halohalo repo
HaloLID().predict("Maayong buntag sa inyong tanan")     # -> ceb

Or with fastText directly:

import fasttext
from huggingface_hub import hf_hub_download
m = fasttext.load_model(hf_hub_download("sapinsapin/halo-lid", "model.ftz"))
m.predict("kumusta ka na kaibigan")      # -> __label__fil

Input should be lowercased with digits and punctuation stripped; see normalize() in halolib/lid.py. For documents, prefer predict_document(), which votes over sentences weighted by length and also returns an agreement score β€” useful for spotting code-switched pages.

Licence and limitations

cc-by-nc-4.0, because part of the training text comes from PLD, which is CC-BY-NC and research-only. Treat this model as research-use.

  • other covers only Indonesian and Malay. Spanish, Chavacano or other Philippine languages outside the ten will be forced into the nearest label, usually with low confidence.
  • Trained mostly on read prompts and web prose; performance on very short, noisy, or heavily code-switched text (Taglish) is lower β€” the <=5 words column is the honest indicator.
  • fil covers Tagalog and Filipino together; PLD does not distinguish them.
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support