lance-a-laya: football highlights from Brazilian commentary

This is a fine-tune of Laya, an open decision model. It reads 20 seconds of Brazilian Portuguese football commentary and answers five questions in one forward pass:

question answer
goal: is a goal scored live in this window? yes / no
big_chance: a clear chance that doesn't go in? yes / no
controversy: is a refereeing decision contested or reviewed? yes / no
card: is a yellow or red card shown? yes / no
intensity: how intense is the moment? calm / building / peak

Run it every 5 seconds on a live transcript to flag highlights for clips and alerts. It doesn't generate text: every answer is a probability over the allowed options. The questions and their yes/no criteria are in questions.json.

The code, the data pipeline and the full evaluation are on GitHub: lucasrabay/lance-a-laya.

Results

The test set is the 2022 World Cup final (Argentina–France, Brazilian TV commentary). It was held out from training and labeled by a human.

Goals, whole match. An alert counts if it lands within ±30 s of a real goal.

method goals caught alerts
Keyword rule ("gol" twice in 20 s) 3 of 6 17
Keyword rule, a single "gol" 6 of 6 38
Laya before fine-tuning 5 of 6 33
This model 6 of 6 6

Window level. 141 human-labeled windows; 9 marked unsure are excluded. F1 uses thresholds chosen on validation; threshold-free average precision is in parentheses.

method goal big chance controversy card intensity acc.
Keyword rules 0.17 (0.27) 0.40 (0.59) 0.47 (0.43) 0.29 (0.40) 0.38
Laya before fine-tuning 0.29 (0.31) 0.08 (0.10) 0.24 (0.26) 0.06 (0.08) 0.42
This model 0.70 (0.93) 0.62 (0.66) 0.67 (0.78) 0.57 (0.64) 0.72
LLM teacher (Claude), for reference 0.87 (0.85) 0.77 (0.84) 0.57 (0.46) 0.80 (0.52) 0.79

Calibration. Pooling the 4 yes/no questions on the human-labeled set, ECE is 0.037 with calibration.json and 0.056 without it.

Speed

setup p50 per window, all 5 questions
This model, NVIDIA T4 (batch 1) 31.6 ms (46.2 windows/s batched)
This model, NVIDIA L4 (batch 1) 37.8 ms (58.7 windows/s batched)
Qwen2.5-3B-Instruct-AWQ on vLLM, same T4 type, writing the same 26-token JSON answer 366.1 ms
  • How the LLM was timed: both rows on the T4 were timed by the author, on the same 200 validation windows, in process, at batch 1. The LLM row used prefix caching.
  • Over the network: from a laptop in Brazil to a cloud T4 and back, the p50 is 262 ms.

How to use

# pip install laya==0.3.24 huggingface_hub
import json
import laya
from huggingface_hub import snapshot_download

path = snapshot_download("LucasRabay/lance-a-laya")
agent = laya.load(path, device="cuda", calibration=f"{path}/calibration.json")  # or device="cpu"
questions = json.load(open(f"{path}/questions.json"))
thresholds = json.load(open(f"{path}/thresholds.json"))["thresholds"]

state = {"narracao": "Lá vai ele, ajeitou pra bater, bateu... é gol! Gol do Brasil!"}  # invented example
answers = agent.predict(state, questions)["answers"]

flags = {q: answers[q]["noul"] >= t for q, t in thresholds.items()}  # noul = P(yes)
levels = ["calm", "building", "peak"]
p = answers["intensity"]["probabilities"]  # {level index: probability}
intensity = levels[int(max(p, key=p.get))]
print(flags, intensity)
  • State: the state is {"narracao": <transcript of the last 20 s>}.
  • Transcripts: the GitHub repo transcribes with faster-whisper large-v3, with VAD turned off. VAD drops goal calls shouted over crowd noise.
  • Validation sets: thresholds.json and calibration.json were fit on the two validation matches.

Training

  • Base: convaiinnovations/laya-multilingual (laya 0.3.24), whose weights are identical to convaiinnovations/laya multilingual/.
  • Matches:
    • Training: 7 matches, five from the 2022 World Cup with CazéTV commentary and two from the 2024 Copa América with Rádio Tupi and Rádio Bandeirantes commentary.
    • Validation: 2 matches.
    • Test: the 2022 final.
  • Windows: 20 s windows every 5 s.
  • Labels:
    • Goals and cards come from StatsBomb open-data events aligned to the audio.
    • Big chance and controversy combine StatsBomb shots and penalty fouls with spans from LLM (Claude) labelers that read the transcripts.
    • Intensity comes from the LLM labelers alone.
    • Spans the labelers marked uncertain were left out of training.
    • Only the human-labeled test set is used for evaluation.
  • Run: Laya's official single-device fine-tuning script, on 16,027 question items after rebalancing: 3 epochs, max_len 512, micro-batch 8, 830.8 s on one A100.
  • Calibration: temperatures refit on the validation matches. The yes/no temperature is 4.65; the score temperature sits at the runtime clamp of 5.0.

Limitations

  • Narrow scope: PT-BR football commentary only. It was tested on a single held-out match with a single labeler. It is not a general classifier.
  • The LLM teacher is still stronger on window F1 for goal, big chance, card and intensity.
  • Uneven results by question:
    • Cards tie with the keyword rule at event level, because narrators say "cartão amarelo" literally.
    • Quiet referee disputes are missed.
  • Thresholds are brittle: calibrated probabilities top out near 0.98, and several misses sit 0.001–0.006 below a threshold.
  • No abstention: the action head always acts (act_probability was 1.0 on every answer), so don't rely on it to escalate.
  • Not tested on a live broadcast: it is built for live use, but it was tested on transcripts replayed in real time.
  • No decisions about people: don't use it for anything outside sports highlights.

Data and copyright

The training transcripts come from copyrighted broadcasts and are not redistributed here. The GitHub repo has the window IDs, labels, predictions and evaluation code; it has no audio, video or transcript text.

License

Apache-2.0, the same as the base model. Laya is by Convai Innovations.

Downloads last month
12
Safetensors
Model size
0.3B params
Tensor type
F16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LucasRabay/lance-a-laya

Finetuned
(70)
this model