Fix model_type in config_sentence_transformers.json, add Sentence Transformers usage

#1
by tomaarsen HF Staff - opened

Hello!

I recently released the new MultiVectorEncoder class in Sentence Transformers v6.0, and colbert-ko-0.1b is a great fit for it. You can read more about the release here: https://huggingface.co/blog/multi-vector-encoder.

Heads up, this rest of this PR was AI-generated and human-reviewed.

Pull Request overview

  • Fix model_type in config_sentence_transformers.json, which currently disables the [Q] /[D] markers and the query expansion for any Sentence Transformers load.
  • Add a Sentence Transformers usage section to the model card, plus the multi-vector tag.

Details

config_sentence_transformers.json declares model_type: "ColBERTWrapper", which the loader does not recognise, so the load falls off the PyLate translation path and silently drops the [Q] /[D] prefix tokens and the 32-token query expansion. Queries come out as 12 raw tokens instead of 32, and the documents land at a per-token cosine of about -0.1 against the reference. Setting it to "ColBERT", the value PyLate itself writes, restores all of it, and the weights are untouched.

Verified against this repository's own bundled model.py path on transformers 5.15, which is the reference your card documents: per-token cosine 1.000000 on the query and both documents, maximum absolute deviation around 1e-06, and MaxSim scores identical to four decimals (29.6262 and 21.1225 from both). The module pipeline Sentence Transformers builds, Transformer -> Dense -> MultiVectorMask -> Normalize, mirrors model.py step for step, with 1_Dense/ supplying the 128-dim Matryoshka head exactly as your PyLate note describes.

pip install "sentence-transformers>=6.0.0"
from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder("dragonkue/colbert-ko-0.1b")

query = "λŒ€ν•œλ―Όκ΅­μ˜ μˆ˜λ„λŠ” μ–΄λ””μΈκ°€μš”?"
documents = [
    "λŒ€ν•œλ―Όκ΅­μ˜ μˆ˜λ„λŠ” μ„œμšΈμ΄λ©°, 인ꡬ가 κ°€μž₯ λ§Žμ€ λ„μ‹œμ΄λ‹€.",
    "νŒŒλ¦¬λŠ” ν”„λž‘μŠ€μ˜ μˆ˜λ„μ΄κ³  μ—νŽ νƒ‘μœΌλ‘œ 유λͺ…ν•˜λ‹€.",
]

query_embeddings = model.encode_query(query)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings[0].shape)
# torch.Size([32, 128]) torch.Size([18, 128])

# MaxSim late-interaction scoring (higher is more relevant)
scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
# tensor([[29.6262, 21.1225]], device='cuda:0')

To try this before merging, pass revision="refs/pr/1" to MultiVectorEncoder.

Happy to tweak anything you'd like changed. Please let me know if you have any questions or feedback!

  • Tom Aarsen
tomaarsen changed pull request status to open
Ready to merge
This branch is ready to get merged automatically.

Sign up or log in to comment