Skip to content

Trust scores

The measured accuracy of Cogeto, per release

Current release v1.9.02026-08-20

Cogeto publishes its own measured accuracy for every release, the same way a service publishes uptime, including the numbers that fall short of their targets. The numbers are read straight from the per-release data files at build time, including the ones that moved the wrong way.

Model configuration
Language
Compare

Aggregate blends the per-language corpora. It is shown so a weak language can never hide inside an average: switch the selector to read each language on its own.

Every gate floor is set at the honest current value of the metric, never at a target the project has not reached, and floors only ratchet upward. Floors apply per language as well as in aggregate, so the gate you see here changes with the language you select.

Current scores

Extraction and reconciliation quality for the selected model configuration and language, measured against a hand-labeled golden corpus.

Scores for configuration mistral-default, Aggregate, v1.9.0
MetricAggregate
83.6%
95.8%
92.9%
94.4%
80%
85.7%
66.7%9 pairs
90.6%

Chat suite

End-to-end question-and-answer cases. A pass means the answer was grounded in the right facts from the corpus. Failing case ids are published, not hidden.

32/35cases pass

Failing case ids: changed_since, strict_mode_hr, who_is_ana

Trends

The ten most recent releases from the v1 line on, oldest to newest, on an honest 0 to 100 percent axis. The complete history is kept, and every release stays immutable once published. The dashed line is the continuous-integration gate that a release must clear to ship.

Measured at releaseBackfilled: transcribed from recorded runs rather than emitted by the harness at release time.CI gate
Extraction precision83.6%
Extraction recall95.8%
Verification agreement92.9%
Deduplication accuracy94.4%
Contradiction precision80%
Contradiction recall85.7%
Supersedes accuracy66.7%
Query-rewrite routing accuracy90.6%
Trend data for configuration mistral-default, Aggregate, all releases
ReleaseExtraction precisionExtraction recallVerification agreementDeduplication accuracyContradiction precisionContradiction recallSupersedes accuracyQuery-rewrite routing accuracy
v1.2.079.2%91.3%89.2%92.9%not measured100%not measurednot measured
v1.3.079.6%92%86.9%92.9%not measured100%not measurednot measured
v1.4.081.3%93.8%92.9%92.9%not measured100%not measurednot measured
v1.4.178.9%91.1%83.3%92.9%66.7%100%75%90.6%
v1.4.2-local78.9%91.1%83.3%92.9%66.7%100%75%90.6%
v1.5.082.7%93.9%92.6%94.4%75%100%75%90.6%
v1.6.082.3%90.4%91.7%94.4%82.4%100%66.7%84.4%
v1.7.184.2%95.8%88.9%94.4%81.3%92.9%66.7%90.6%
v1.8.084.1%94.6%91.9%94.4%81.3%92.9%66.7%90.6%
v1.9.083.6%95.8%92.9%94.4%80%85.7%66.7%90.6%

Notes from the releases

  • v1.4.2-local
    • Self-hosted configuration measurement (llama.cpp ff711 + bge-m3 on the operator's own hardware), thinking suppressed on every call. Advisory against the Mistral-measured gates: contradiction recall and query-rewrite accuracy sit below their floors; both zero-tolerance gates (injection, subject) pass. Not a release measurement.

Provenance

Each release, with the exact commit it was measured at, the harness version, and the corpus sizes. Published files are never edited after release, so a number here can always be traced back to the tree it was measured on.

  • v1.9.02026-08-201c6a5990c0
    Harness
    extraction/v0006 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0002 · query_rewrite/v0006 · thresholds v1 + chat answer/v0009 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 102 golden cases (50 english, 52 croatian) · 54 reconciliation pairs · 35 chat cases

  • v1.8.02026-08-164e9e590862
    Harness
    extraction/v0006 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0002 · query_rewrite/v0006 · thresholds v1 + chat answer/v0009 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 102 golden cases (50 english, 52 croatian) · 54 reconciliation pairs · 35 chat cases

  • v1.7.12026-08-1467f8e10bba
    Harness
    extraction/v0006 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0002 · query_rewrite/v0006 · thresholds v1 + chat answer/v0009 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 102 golden cases (50 english, 52 croatian) · 54 reconciliation pairs · 35 chat cases

  • v1.6.02026-08-12ffa3e4b27b
    Harness
    extraction/v0006 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0002 · query_rewrite/v0006 · thresholds v1 + chat answer/v0009 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 102 golden cases (50 english, 52 croatian) · 54 reconciliation pairs · 35 chat cases

  • v1.5.02026-08-050c5f8cc1a2
    Harness
    extraction/v0005 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · query_rewrite/v0006 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 98 golden cases (48 english, 50 croatian) · 33 reconciliation pairs · 24 chat cases

  • v1.4.2-local2026-08-05be2254dc6e
    Harness
    extraction/v0005 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · query_rewrite/v0006 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
    Configuration cogeto-offline

    Models: pipeline cogeto, answer cogeto, embedding bge-m3

    Corpus: 98 golden cases (48 english, 50 croatian) · 33 reconciliation pairs · 24 chat cases

    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 86 golden cases (42 english, 44 croatian) · 29 reconciliation pairs · 24 chat cases

  • v1.4.12026-08-0130d97a1007
    Harness
    extraction/v0004 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · query_rewrite/v0006 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 86 golden cases (42 english, 44 croatian) · 29 reconciliation pairs · 24 chat cases

  • v1.4.02026-07-31eca1222b3b
    Harness
    extraction/v0004 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 86 golden cases (42 english, 44 croatian) · 20 reconciliation pairs · 24 chat cases

  • v1.3.02026-07-305bf26124c0
    Harness
    extraction/v0004 + verification/v0006 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 86 golden cases (42 english, 44 croatian) · 20 reconciliation pairs · 24 chat cases

  • v1.2.02026-07-29b423be0131
    Harness
    extraction/v0002 + verification/v0004 · reconcile_dedup/v0001 + reconcile_contradiction/v0001 · thresholds v1 + chat answer/v0007 · grader eval-coverage/v0001
    Configuration mistral-default

    Models: pipeline mistral-small-latest, answer mistral-medium-latest, embedding mistral-embed

    Corpus: 76 golden cases (37 english, 39 croatian) · 20 reconciliation pairs · 24 chat cases

Back to cogeto.eu