Access to memory-resoner

Released for research, evaluation, internal validation and education. Tell us who you are and we will grant access.

By requesting access you agree to the Elda Community License 1.0: no commercial use, no redistribution of the weights or derivatives, and attribution as "Built with Elda". Commercial licensing is available on request.

Log in or Sign Up to review the conditions and access this model content.

memory-resoner β€” the reference step of conversational memory

A conversation can only be stored if you know what γ€Œκ·Έ 동넀」 meant. memory-resoner reads the turns of a conversation and says, for each mention, which earlier mention it refers to β€” so the layer above can write a memory entry that points at the right words.

"합정동에 κ°€κ²Œλ₯Ό μ—΄μ–΄μš” . κ·Έ λ™λ„€λŠ” μœ λ™μΈκ΅¬κ°€ λ§Žλ‚˜μš” ?"

  mentions   합정동에   (s0, w0–0)
             κ·Έ λ™λ„€λŠ”  (s1, w0–1)
  chains     [합정동에 Β· κ·Έ λ™λ„€λŠ”]

version 0.2.0 Β· 308M parameters Β· fp32 Β· 33 ms per sentence on CPU when the mentions are supplied (125 ms when it finds them itself).


What changed in 0.2.0 β€” it resolves reference in more than one language now

0.1.4 was a Korean model. Asked to resolve a reference in Japanese it scored 0.0000 β€” not "poorly", but nothing: it could not find a single mention in a language written without spaces. 0.2.0 is the version where that stops being true.

Japanese reference resolution      0.0000  β†’  **0.5664**
CorefUD, 19 languages, end-to-end  0.0699  β†’  **0.1842** macro-F1   (17 of the 19 never seen in training)

(An earlier revision of this card listed a third line here β€” "longest input 512 β†’ 8,192". That line was wrong in both directions and has been withdrawn; see Corrections at the end.)

Three things had to change together, and none of them alone would have done it.

β‘  The encoder. The old one was Korean-specialised β€” 149M parameters, a 50k vocabulary. The new one is 308M with a 256k multilingual vocabulary. Both are ModernBERT; what changed is the language coverage, not the context length (this bundle encodes max_len=384 either way β€” see Corrections). The heads above it, the output contract and every entry point are unchanged; you call this model exactly as before.

β‘‘ The training mix. Korean CorefUD plus in-house synthetic dialogue became Korean and English CorefUD plus in-house synthetic dialogue. The other 17 languages were never trained on β€” what they get is transfer.

β‘’ β˜… How the surface is rebuilt β€” and this one is easy to get wrong. This model reads a list of words and rebuilds the sentence to feed the encoder. Rebuilding with a space between every word is correct for Korean and English and wrong for Japanese, Chinese and Thai, whose "words" come out of a segmenter and were never separated in the first place. Put spaces back in and you hand the encoder a surface no one has ever written:

Japanese reference resolution score
word_joiner="" (as written) 0.5664
word_joiner=" " (words spaced out) 0.2301

Same weights, same inputs, same knob positions otherwise β€” far apart. The knob defaults to " "; pass "" for languages written without spaces. A bigger encoder would not have saved you from this.

⚠ How much text this model actually reads: 384 tokens. That is what config.max_len ships as, and it is what the weights were trained with β€” it has not changed between versions. The backbone could hold more, but the bundle does not use it, and raising the flag today would feed the heads a regime they never saw. Reading longer context is a future version's job.

What you can check yourself

0.2.0 reproducible here?
Japanese reference resolution † 0.5664 βœ—
CorefUD, 19 languages, macro end-to-end 0.1842 βœ“ public data
in-domain span agreement (exact) ‑ 97.5% βœ—
input positions this bundle uses 384 βœ“ config.max_len

† JMultiWOZ 1.0 (CC BY-SA 4.0), references mined from that corpus's own slot annotation, words segmented with janome, 113 in-window cases, mentions supplied. ‑ measured on an in-house holdout of machine-generated dialogue, which we do not redistribute, so you cannot re-run it. It is listed because it is the axis our own consumers read, not as a claim you should take on trust. Rows marked βœ“ you can check yourself with the two scripts in this repository.

New knobs (all optional, all defaulting to the old behaviour)

field default what it does
detect_layers 1 BIO layers in the detection head. 1 emits outermost mentions only, as before.
word_joiner " " What goes between words when the surface is rebuilt. Pass "" for languages written without spaces (Japanese, Chinese, Thai) β€” otherwise the encoder is handed a surface it has never seen.
null_bias 1.0 How strongly to discourage the "no antecedent" choice. Larger means the model answers more often. Raise it only on a path where every mention really has an antecedent β€” see the warning below.

Each can also be passed per call: model.resolve_last(sentences, mentions, tok, joiner="", null_bias=2.0).

β›” Do not set fix_mistral_regex=True

Recent transformers prints a warning when loading this tokenizer, suggesting the flag. We measured it on 5 sentences (ko, ja, en): with the flag, every one of them tokenizes to byte garbage β€” 강남역 becomes Γͺ Β° Δ· Γ« Δ€ Β¨ instead of ▁강 남 μ—­, and sequence length goes from 19 to 71. The model still runs and still returns spans; they are simply wrong. The default path is the correct one, and it is the one every number on this card was measured with.

⚠ null_bias: the number that looks free on the wrong test set

On a benchmark that only ever asks "which earlier mention does this pronoun refer to?", every abstention is wrong by construction, so pushing the model to always answer looks strictly better β€” and keeps looking better the harder you push. On text where "no antecedent" is sometimes the correct answer, the same push is destructive. We measured both:

null_bias KLUE-WoS, given mentions (no NULL-correct cases) CorefUD ko (NULL-correct cases present)
0 0.4409 0.5588
1 (shipped) 0.5074 0.5714
3 0.5443 0.5551
10 0.5616 0.1670

1.0 is the only setting that improves both, and it does not trade away precision (accuracy among answered cases: 0.5576 β†’ 0.5583). Going higher buys one number by destroying another. If your pipeline guarantees that the mentions you pass in are anaphoric, a higher value is safe; otherwise leave it alone.

Where it sits

memory-resoner is one stage of a conversational stack β€” Korean first, and from 0.2.0 also English and Japanese. Each stage answers one question and hands the next stage a span, never prose.

utterance
   β”‚
   β”œβ”€ Elda-AI/intenter              what was said, and what kind of thing it is  ── spans come from here
   β”œβ”€ slot extraction               which of those the user said about themselves
   β”‚
   β–Ό
β˜… memory-resoner                    β˜… which earlier mention does this one refer to
   β”‚
   β–Ό
memory write                        a pointer into the user's own words β€” checkable, not generated

Where the spans come from. In the Elda stack, mention spans are produced upstream by Elda-AI/intenter and this model consumes them; its card is the reference for how a span is drawn, which types exist, and what the channels mean. Used standalone, memory-resoner will find its own mentions instead.

⚠ Those two paths are not equally good, and the gap is large. On KLUE-WoS dialogue reference, this version scores 0.5074 when the mentions are supplied and 0.2547 when it has to find them itself. Detection and linking are two abilities and the end-to-end number is their product β€” so if you already have spans upstream, pass them in.

⚠ The two sides count spans in different units, and the conversion is yours to make. Upstream spans are character offsets with the particle left outside (γ€Œν•©μ •λ™γ€). This model works in μ–΄μ ˆ (word) indices, and a μ–΄μ ˆ contains its particle, so the same mention is γ€Œν•©μ •λ™μ—γ€ here. Neither is wrong; they are different units. Align on the μ–΄μ ˆ that contains the upstream span.

What it returns. Span indices and chains. Not a rewritten sentence, not a summary β€” an index into the words you sent, which the layer above can verify before storing anything.

What it does not decide. Whether a fact is true, whether it is worth keeping, how long it lives. Those belong to the system around it. This stage answers one question and stops.


Why a small encoder

this model LLM prompting
Parameters 308M 7B – 70B
Latency per sentence, CPU β€” linking, mentions given p50 33 ms seconds
Latency per sentence, CPU β€” finding mentions and linking p50 125 ms seconds
Output span index β€” checkable byte-for-byte free text to be parsed
Determinism same input, same chains sampling

Memory is written on every turn. A stage that runs that often has to be cheap, and its answer has to be something the next stage can check rather than trust. A span index is both.


Output contract

input        sentences, in order β€” a string per sentence, or a list of μ–΄μ ˆ
output       mentions Β· antecedents Β· chains (links closed transitively)
spans        inclusive word (μ–΄μ ˆ) indices β€” Korean particles stay attached (see the note above)
window       6 previous sentences of context Β· 40 previous mentions as candidates

A mention with no antecedent opens a chain of its own β€” that is how the next stage learns it has seen a new entity, so singleton chains are kept rather than dropped.


Usage

from transformers import AutoModel, AutoTokenizer

m   = AutoModel.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True).eval()
tok = AutoTokenizer.from_pretrained("Elda-AI/memory-resoner", trust_remote_code=True)

out = m.coref(["합정동에 κ°€κ²Œλ₯Ό μ—΄μ–΄μš” .",
               "κ·Έ λ™λ„€λŠ” μœ λ™μΈκ΅¬κ°€ λ§Žλ‚˜μš” ?"], tok)

out["mentions"]     # [{'sent': 0, 'words': [0, 0], 'text': '합정동에'}, ...]
out["antecedents"]  # [None, 0]   ← per mention: what it points back at
out["chains"]       # [[0, 1]]
call does
m.coref(sentences, tok) finds the mentions and links them
m.mentions(sentences, tok) finds mentions only β€” a span provider for another linker
m.resolve(units, mentions, tok) links only β€” you supply the mentions
m.coref_last(sentences, tok) β˜… serving shape β€” only the newest turn's mentions are returned, with the earlier turns kept as context
m.resolve_last(sentences, mentions, tok) β˜… serving shape, mentions supplied β€” the same, but you hand in the spans

The two _last calls are what a live conversation wants: one turn arrives, you want answers for that turn, and the turns before it are context rather than output. They take the whole conversation and cut the window themselves, so the caller never has to know how long the window is.

β›” On coref_last, do not count with antecedents

The two calls differ here, and it matters. resolve_last returns all the mentions you gave it, so its antecedents[i] indexes that same list and reads normally. coref_last returns only the newest turn's mentions β€” so when the antecedent sits in an earlier turn, which on a live conversation is almost always, antecedents[i] is None. There, None wears one face for two different facts: "it pointed at something outside this list" and "it pointed at nothing", and a consumer that counts non-None gets a structural zero even when the model is working.

field what it is use it for
linked[i] β˜… bool β€” did this mention resolve at all counting, on either call
n_linked sum of the above counting
antecedent_mentions[i] the antecedent itself ({sent, words, text}), or None reading the answer
antecedents[i] index into the returned list, or None safe on resolve_last; on coref_last only for within-turn links

resolve_last also reports what it had to drop or fix, and none of it is thrown away silently: mentions_outside_window Β· mentions_unencodable Β· window_truncated Β· mentions_out_of_order. ⚠ When mentions_out_of_order > 0 the layer re-sorted your mentions into reading order, and then antecedents[i] > i can occur β€” the index is into the list you gave, not into reading order. If that count is zero, "an antecedent index is always smaller than its own" holds.

Notes that matter in practice:

  • fp32. The config pins it. In bf16 near-ties flip and chains change.
  • Pass sentences, not a paragraph. The sentence boundary is what the window is counted in.
  • Send the resolved span downstream, not the reference. The output is an index into the words you sent, so whatever consumes it can verify what it was handed.
  • Strip the particle at display time, not at span time. Keeping it inside the μ–΄μ ˆ is what makes the span a plain index into your own input; trimming belongs to whatever renders the value.

Measurements

Public Korean coreference, held out from training; document overlap with the training split is 0 and the benchmark script asserts it. Both scripts ship in this repository.

⚠ If you are replacing 0.1.4 in a Korean-only pipeline, measure before you swap. The Korean scores below sit under 0.1.4's β€” this release spent Korean headroom to reach other languages and longer inputs. The two scripts here let you check that against your own data rather than take our word for it.

pip install torch transformers datasets
python bench_corefud_ko.py --model Elda-AI/memory-resoner

Mention detection β€” boundaries must match exactly

precision recall F1
0.8354 0.6145 0.7081

Linking, mentions given

scored mentions 10263
accuracy 0.8305
non-NULL accuracy 0.4616 (n=2721)
deictic, non-NULL 0.4207 (n=271)
reference β€” always NULL 0.7349
reference β€” always most-recent 0.0448

Read the non-NULL rows. Most mentions open a chain, so answering NULL every time already scores 0.7349; the rows that require actually linking are the ones in bold.

End-to-end β€” (mention, antecedent) pairs, the model finding its own mentions

precision recall F1
0.6554 0.508 0.5724

Every run first pushes the gold answers back through the scorer and asserts they survive intact (1.0000). A scorer that cannot return its own gold is not grading anything.

On dialogue

The numbers above are prose. On dialogue β€” references mined from a public multi-turn Korean corpus (KLUE-WoS), with the antecedent taken from that corpus's own annotation rather than chosen by us β€” this version resolves γ€Œκ·Έ 식당」-style references at 0.5074 when the mentions are supplied, and 0.2547 when it has to find them itself.

python bench_wos_references.py --model Elda-AI/memory-resoner

β˜… That difference is the whole point of the two paths. End-to-end is detection times linking, so a pipeline that already has spans should pass them in rather than ask for them again. ⚠ These are measured with candidates mined by the gold rule, which makes them a perfect-upstream ceiling β€” a real intenter will score lower. If your text is prose rather than conversation, 0.1.1 is the stronger checkpoint and it remains in this repository's history.


What comes next

Conversational memory here is built in three layers. They are not future versions of this model β€” they are different kinds of component, and saying so is the point.

what it is for how it is built
Layer 1 Β· conversation memory what was said in this conversation, and what refers to what β˜… this model + deterministic assembly
Layer 2 blackbox β€” how a memory is addressed and kept apart from every other conversation not described here
Layer 3 Β· authority and persona whether something holds and on what grounds, and what an agent is like across conversations a deterministic VM, and a decoder used off the request path

Layer 1 is the only one this model sits in. Layer 2 is deliberately not described: it is where memories are separated from one another, and it is the part of the system we keep closed. Layer 3 is where facts acquire grounds and where a persona is formed, and it is built from two very different machines β€” one that must be exact, and one that must be fluent.

The VM β€” the exact one. Memory is written in a small closed language and executed by a deterministic machine, so a write is either accepted, a duplicate, a replacement, or refused, and the same inputs always produce the same store. Reads are queries against that store, not recollection. This is what makes memory auditable rather than merely plausible, and it is why the per-turn path carries no generation at all.

The decoder β€” the fluent one. Never on the request path. Writing and reading memory happen on every turn and stay on the encoder-and-machine side. Generation is for prose a person will read, and for periodic passes that look across many conversations at once to form a persona β€” a pass that can fail without the conversation stopping.


Versioning

v<major>.<minor>.<patch>, three digits. Minor moves when the task, the inputs or the heads change β€” when your code has to change. Patch moves when the same job is done better. The version lives in config.json and is cross-checked against this card at publish time.

Current: 0.2.0. Earlier weights remain in this repository's history.

β˜… 0.2.0 is a baseline reset. The encoder underneath changed, so its numbers are not comparable to earlier versions on every axis β€” some went up (other languages, long inputs), some went down (Korean dialogue reference). Treat this version, not the earlier ones, as the reference point for judging what comes next.

This repository previously hosted a different model, a Korean slot tagger. It is not an earlier version of this one.


License & access

Released under the Elda Community License 1.0 (see LICENSE). Provenance of the training data is recorded in LINEAGE.md; no user conversations were used.

β›” Correction to the data lineage published for 0.1.4

The LINEAGE.md shipped with 0.1.4 stated that no in-house synthetic data was used. That was wrong. 0.1.4 was trained on 14,848 in-house synthetic dialogue units (machine generated from templates) in addition to the public CorefUD gold, and those units did carry loss β€” on the linking head only, never on detection. The claim was hand-written into the file rather than derived from the training run, and nobody re-read it when the recipe changed.

We are correcting it rather than quietly replacing it. From 0.2.0 on, LINEAGE.md is generated from the training log and the packaging tool refuses to assert provenance without one. The statement that no user conversations were used was true then and remains true.

β›” Corrections to the first published revision of this card (0.2.0)

Two numbers in the first upload were wrong. Both are corrected above; we record what they were and why, rather than replacing them silently.

β‘  The Japanese score was measured with a different knob setting than the one shipped. The card said 0.3097 (and 0.0619 for the spaced variant). Those came from an internal scoring bundle with null_bias=0, while this bundle ships null_bias=1.0. Re-measured on the shipped bundle with the same probe, the same 113 in-window cases and the same arguments, the numbers are 0.5664 and 0.2301 β€” the card understated its own model. A score is not a property of the weights alone; it is a property of the weights and the knobs, which is why every number in this card now names the knob positions it was taken at, and the tool that fills this card refuses a measurement whose knobs disagree with config.json.

β‘‘ "Longest input it can read whole: 512 β†’ 8,192" was wrong in both directions. The previous encoder's config declares 16,384 positions, not 512 β€” read by the same field that gave 8,192 for this one β€” so the swap reduced the backbone's positional capacity rather than raising it. And neither number describes this bundle, which encodes max_len=384 and always has, in this version and the previous one. The line has been withdrawn: there was no gain on that axis, and nothing was lost in what the model actually reads.

  • βœ… Research, evaluation, internal validation, education β€” free of charge
  • ✳ Attribution: "Built with Elda"
  • β›” Commercial use and redistribution require a separate agreement

Access is gated: tell us who you are and access is granted automatically.

Citation

@software{memoryresoner2026,
  title  = {memory-resoner: reference resolution for Korean conversational memory},
  author = {Elda AI},
  year   = {2026},
  url    = {https://huggingface.co/Elda-AI/memory-resoner}
}
Downloads last month
4
Safetensors
Model size
0.3B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support