feat(parties): selection-step evaluation against document-anchored truth (#2169)
* feat(parties): selection-step evaluation against document-anchored truth Scores which similar keys are the same party without human labels: every key in the set carries an org number OCR-read from a linked invoice, so two keys are the same party exactly when the org numbers agree. Rules and the Bedrock model both land at 0.91 pair precision; the residual false merges are different legal entities sharing a trade name, which text cannot and should not separate. Results and caveats in the README. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * chore(parties): make the model selector opt-in in the selection eval Sending voucher key text to the AI provider now requires --llm; the default run scores the rules selector only and makes no network call. Documents that the opt-in path uses the same configured provider the production categorizer already sends the same text to. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> --------- Co-authored-by: Jakob Wennberg <311770904+jakobwennberg-oss@users.noreply.github.com> Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Fable 5.1
Jakob Wennberg
parent
5291806c37
commit
69d3bba587
@@ -96,6 +96,38 @@ the default for that reason. The model's remaining party misses are text
|
||||
with a person's name or a card descriptor, exactly where a hard key from a
|
||||
document would decide instead.
|
||||
|
||||
## Selection-step evaluation against document-anchored truth, 2026-09-02
|
||||
|
||||
`eval-selection.ts` scores the selection step (which similar keys in the same
|
||||
company are the same party) without human labels: every key in the set has an
|
||||
org number read by OCR from an invoice linked to its voucher, so two keys are
|
||||
the same party exactly when the org numbers agree. Keys mapping to several
|
||||
org numbers are excluded. 225 anchors, 1,038 candidate pairs (747 same, 291
|
||||
different), 44 anchors with no true match. Draw: `draw-selection-gold.sql`.
|
||||
The model selector is opt-in (`--llm`): it sends the same voucher text the
|
||||
production categorizer already sends to the configured AI provider (Bedrock EU
|
||||
on hosted), so run it only with an env file whose AI settings you have checked.
|
||||
|
||||
| Selector | Pair precision | Pair recall | Non-match recall | Anchors fully right | "None" precision / recall |
|
||||
|---|---|---|---|---|---|
|
||||
| Rules (same core name) | 0.909 | 0.776 | 0.801 | 0.547 | 0.40 / 0.86 |
|
||||
| Sonnet 5 on Bedrock EU, zero-shot | 0.910 | 0.904 | 0.770 | 0.756 | 0.68 / 0.77 |
|
||||
|
||||
What the false merges are: both selectors merge "Levfakt Fortnox (325)" with
|
||||
"Levfakt Fortnox (154)", "Vattenfall Kundservice" with "Vattenfall
|
||||
Eldistribution", "Jämtkraft" with "Jämtkraft" under another org number. Those
|
||||
are different legal entities sharing a trade name, or OCR variance in the org
|
||||
number. Text cannot separate them and should not try: the hard key from the
|
||||
document is the referee, which is exactly the resolver's precedence order.
|
||||
What the model catches that the rules miss: "lån themax" versus "lån från
|
||||
themax", and paraphrases with a changed word order.
|
||||
|
||||
Read with one caveat: sales-side vouchers ("kundbet ...") contribute noisy
|
||||
truth because the org number on the linked document can be the customer's;
|
||||
the next draw restricts to expense-side vouchers.
|
||||
|
||||
Cost: about 1,300 input tokens per anchor, 298k tokens for the run.
|
||||
|
||||
## Still to confirm in this phase
|
||||
|
||||
- SCB access: the free Företagsregistret API carries F-skattstatus,
|
||||
|
||||
@@ -0,0 +1,44 @@
|
||||
-- Parties, phase 0/1: draw the document-anchored ground truth for the
|
||||
-- SELECTION step. Read-only. Output contains customer voucher text: keep it
|
||||
-- in dev_docs/parties/golden/, never in this public repository.
|
||||
--
|
||||
-- Truth without a human: a key whose voucher is linked to a document with an
|
||||
-- OCR-read supplier org number is a known party; two keys in the same company
|
||||
-- are the same party exactly when their org numbers agree. Keys that map to
|
||||
-- several org numbers are dropped as ambiguous; documents whose "supplier"
|
||||
-- org number is the company's own are dropped as the company's sales side.
|
||||
WITH real_co AS (SELECT id, regexp_replace(coalesce(org_number,''), '[^0-9]', '', 'g') AS own_org FROM companies WHERE name NOT ILIKE '%sandl%'),
|
||||
linked AS (
|
||||
SELECT d.company_id, je.description,
|
||||
regexp_replace(coalesce(d.extracted_data->'supplier'->>'orgNumber',''), '[^0-9]', '', 'g') AS org,
|
||||
public.normalize_counterparty_key(je.description) AS k,
|
||||
(SELECT l.account_number FROM journal_entry_lines l WHERE l.journal_entry_id = je.id AND l.account_number ~ '^[4-7][0-9]{3}$' ORDER BY l.debit_amount DESC LIMIT 1) AS acct
|
||||
FROM document_attachments d
|
||||
JOIN real_co c ON c.id = d.company_id
|
||||
JOIN journal_entries je ON je.id = d.journal_entry_id AND je.status = 'posted'
|
||||
WHERE d.extracted_data->'supplier'->>'orgNumber' IS NOT NULL
|
||||
),
|
||||
keys AS (
|
||||
SELECT l.company_id, l.org, l.k, count(*) AS n, mode() WITHIN GROUP (ORDER BY l.description) AS example, mode() WITHIN GROUP (ORDER BY l.acct) AS acct,
|
||||
btrim(regexp_replace(regexp_replace(l.k, '^(levfakt|levfkt|lev\.?fakt\.?|leverantörsfaktura från|leverantörsfaktura|levbet\.?|kvitto|faktura|utgift|inköp)\s+', '', 'g'), '\m\d+\M', '', 'g')) AS core
|
||||
FROM linked l JOIN real_co c ON c.id = l.company_id
|
||||
WHERE length(l.org) = 10 AND l.k <> '' AND l.org <> c.own_org
|
||||
GROUP BY l.company_id, l.org, l.k
|
||||
),
|
||||
uniq AS (SELECT company_id, k, min(org) AS org, sum(n) AS n, min(example) AS example, min(acct) AS acct, min(core) AS core FROM keys GROUP BY company_id, k HAVING count(DISTINCT org) = 1),
|
||||
co_ok AS (SELECT company_id FROM uniq GROUP BY 1 HAVING count(DISTINCT org) >= 2 AND count(*) >= 6),
|
||||
anchors AS (SELECT u.*, row_number() OVER (ORDER BY md5(u.company_id::text || u.k)) AS rn FROM uniq u JOIN co_ok USING (company_id) WHERE length(u.core) >= 4),
|
||||
sample AS (SELECT * FROM anchors WHERE rn <= 320),
|
||||
cands AS (
|
||||
SELECT s.rn AS anchor_rn, o.k, o.example, o.n, o.acct, o.org, extensions.similarity(s.core, o.core) AS sim,
|
||||
row_number() OVER (PARTITION BY s.rn ORDER BY extensions.similarity(s.core, o.core) DESC, o.n DESC) AS cr
|
||||
FROM sample s JOIN uniq o ON o.company_id = s.company_id AND o.k <> s.k
|
||||
WHERE extensions.similarity(s.core, o.core) >= 0.25
|
||||
),
|
||||
agg AS (
|
||||
SELECT s.rn, s.k AS anchor_k, s.example AS anchor_example, s.n AS anchor_n, s.acct AS anchor_acct, s.org AS anchor_org,
|
||||
(SELECT json_agg(json_build_object('k', c.k, 'example', c.example, 'n', c.n, 'acct', c.acct, 'org', c.org, 'sim', round(c.sim::numeric, 2)) ORDER BY c.sim DESC)
|
||||
FROM cands c WHERE c.anchor_rn = s.rn AND c.cr <= 6) AS candidates
|
||||
FROM sample s
|
||||
)
|
||||
SELECT * FROM agg WHERE candidates IS NOT NULL ORDER BY rn LIMIT 300;
|
||||
@@ -0,0 +1,226 @@
|
||||
/**
|
||||
* Parties, phase 0/1: shadow evaluation of the SELECTION step against
|
||||
* document-anchored ground truth.
|
||||
*
|
||||
* WHAT: after blocking, the resolver must decide which of at most six
|
||||
* similar keys in the same company are the same party as an anchor key, or
|
||||
* none. Ground truth here needs no human: every key in the set carries an
|
||||
* org number read by OCR from an invoice linked to its voucher, so two keys
|
||||
* are the same party exactly when their org numbers agree. The draw is
|
||||
* `draw-selection-gold.sql`; keys that map to several org numbers are
|
||||
* excluded because they are ambiguous by construction.
|
||||
*
|
||||
* Two selectors are scored on the same anchors:
|
||||
* 1. rules: a candidate is "same" when its core (AP prefix, supplier number
|
||||
* and digit runs stripped) equals the anchor's core;
|
||||
* 2. the model behind getAiService().generateStructured, asked to return the
|
||||
* indices of candidates that are the same organisation, or none.
|
||||
* Metrics are pair-level (same / different) precision and recall, plus the
|
||||
* anchor-level "none" decision, because a resolver that merges where it
|
||||
* should not is the failure mode the July design warned about.
|
||||
*
|
||||
* SAFETY: read-only. Reads a gitignored JSONL, writes a report next to the
|
||||
* input. Never opens a database connection.
|
||||
*
|
||||
* DATA PROCESSING: the model selector is opt-in (--llm). It sends the voucher
|
||||
* key text and account numbers in the gold file to getAiService(), which is
|
||||
* the same configured provider the production categorizer already sends the
|
||||
* same voucher text to (Claude on AWS Bedrock in the EU on hosted). Run it
|
||||
* only with an env file whose AI configuration you have checked; without
|
||||
* --llm the script scores the rules selector alone and makes no network call.
|
||||
*
|
||||
* Usage:
|
||||
* npx tsx scripts/parties/eval-selection.ts \
|
||||
* --gold dev_docs/parties/golden/selection-2026-09-02.jsonl --env .env.local \
|
||||
* [--out <report.json>] [--llm] [--limit 300]
|
||||
*/
|
||||
import { readFileSync, writeFileSync } from 'node:fs'
|
||||
import { resolve } from 'node:path'
|
||||
import { config as dotenv } from 'dotenv'
|
||||
import { z } from 'zod'
|
||||
|
||||
interface Cand {
|
||||
k: string
|
||||
example: string
|
||||
n: number
|
||||
acct: string | null
|
||||
org: string
|
||||
sim: number
|
||||
}
|
||||
interface Anchor {
|
||||
rn: number
|
||||
anchor_k: string
|
||||
anchor_example: string
|
||||
anchor_n: number
|
||||
anchor_acct: string | null
|
||||
anchor_org: string
|
||||
candidates: Cand[]
|
||||
}
|
||||
|
||||
function arg(name: string): string | undefined {
|
||||
const i = process.argv.indexOf(`--${name}`)
|
||||
return i >= 0 ? process.argv[i + 1] : undefined
|
||||
}
|
||||
const flag = (name: string) => process.argv.includes(`--${name}`)
|
||||
|
||||
const goldPath = resolve(arg('gold') ?? 'dev_docs/parties/golden/selection-2026-09-02.jsonl')
|
||||
const envPath = resolve(arg('env') ?? '.env.local')
|
||||
const outPath = resolve(arg('out') ?? goldPath.replace(/\.jsonl$/, '') + '.eval.json')
|
||||
const limit = Number(arg('limit') ?? 300)
|
||||
const runLlm = flag('llm')
|
||||
dotenv({ path: envPath })
|
||||
|
||||
const anchors: Anchor[] = readFileSync(goldPath, 'utf8')
|
||||
.split('\n')
|
||||
.filter(Boolean)
|
||||
.map((l) => JSON.parse(l) as Anchor)
|
||||
.slice(0, limit)
|
||||
|
||||
// ── Rules selector ──────────────────────────────────────────────────────────
|
||||
const AP = /^(levfakt|levfkt|lev\.?fakt\.?|leverantörsfaktura från|leverantörsfaktura|levbet\.?|kvitto|faktura|utgift|inköp)\s+/
|
||||
const SUFFIX = /\b(ab|aktiebolag|hb|kb|sverige|sweden|ltd|limited|oy|gmbh|inc|sarl|publ|filial)\b/g
|
||||
export function core(k: string): string {
|
||||
return k
|
||||
.toLowerCase()
|
||||
.replace(AP, '')
|
||||
.replace(/\b\d+\b/g, '')
|
||||
.replace(SUFFIX, '')
|
||||
.replace(/[^a-zåäöé ]+/g, ' ')
|
||||
.split(/\s+/)
|
||||
.filter(Boolean)
|
||||
.join(' ')
|
||||
}
|
||||
function rulesSelect(a: Anchor): number[] {
|
||||
const ac = core(a.anchor_k)
|
||||
if (!ac) return []
|
||||
return a.candidates.map((c, i) => (core(c.k) === ac ? i : -1)).filter((i) => i >= 0)
|
||||
}
|
||||
|
||||
// ── Model selector ──────────────────────────────────────────────────────────
|
||||
const SYSTEM = `Du avgör identitet mellan motparter i svensk bokföring.
|
||||
Du får ett ankare (en normaliserad verifikationstext från ett bolags bokföring) och upp till sex kandidater ur samma bolag som liknar det.
|
||||
Svara med index för varje kandidat som är SAMMA organisation som ankaret, alltså samma juridiska person som skulle ha samma organisationsnummer. Olika bolag i samma koncern, franchisetagare med samma kedjenamn, eller ett företag och dess inkassobolag är INTE samma.
|
||||
Leverantörsnummer inom parentes, fakturanummer och prefix som "levfakt" eller "leverantörsfaktura från" är brus, inte identitet.
|
||||
Om ingen kandidat är samma organisation, svara med en tom lista. Gissa inte: hellre tom lista än en felaktig sammanslagning.`
|
||||
|
||||
const Out = z.object({ same: z.array(z.number().int().min(1).max(6)) })
|
||||
const jsonSchema = {
|
||||
type: 'object',
|
||||
properties: { same: { type: 'array', items: { type: 'integer', minimum: 1, maximum: 6 } } },
|
||||
required: ['same'],
|
||||
additionalProperties: false,
|
||||
}
|
||||
|
||||
async function modelSelect(a: Anchor): Promise<{ picks: number[]; usage: unknown; model: string }> {
|
||||
const { getAiService } = await import('@/lib/ai')
|
||||
const service = getAiService()
|
||||
const prompt =
|
||||
`Ankare: "${a.anchor_k}" (exempel "${a.anchor_example}", konto ${a.anchor_acct ?? '?'}, ${a.anchor_n} verifikat)\n` +
|
||||
`Kandidater:\n` +
|
||||
a.candidates
|
||||
.map((c, i) => `[${i + 1}] "${c.k}" (exempel "${c.example}", konto ${c.acct ?? '?'}, ${c.n} verifikat, textlikhet ${c.sim})`)
|
||||
.join('\n') +
|
||||
`\nVilka kandidater är samma organisation som ankaret? Svara med index, eller tom lista.`
|
||||
let lastErr: unknown
|
||||
for (let attempt = 0; attempt < 2; attempt++) {
|
||||
try {
|
||||
const r = await service.generateStructured({
|
||||
tier: 'assistant',
|
||||
system: SYSTEM,
|
||||
prompt,
|
||||
maxTokens: 512,
|
||||
schema: { name: 'same_party', description: 'Indices of candidates that are the same organisation', jsonSchema },
|
||||
})
|
||||
const parsed = Out.parse(r.value)
|
||||
const picks = [...new Set(parsed.same.map((i) => i - 1).filter((i) => i >= 0 && i < a.candidates.length))]
|
||||
return { picks, usage: r.usage, model: r.model }
|
||||
} catch (e) {
|
||||
lastErr = e
|
||||
}
|
||||
}
|
||||
throw lastErr
|
||||
}
|
||||
|
||||
// ── Scoring ─────────────────────────────────────────────────────────────────
|
||||
interface Score {
|
||||
anchors: number
|
||||
pairs: number
|
||||
pair_precision: number
|
||||
pair_recall: number
|
||||
pair_tnr: number
|
||||
anchor_exact: number
|
||||
none_precision: number
|
||||
none_recall: number
|
||||
false_merges: { anchor: string; candidate: string; sim: number }[]
|
||||
missed: { anchor: string; candidate: string; sim: number }[]
|
||||
}
|
||||
function score(picksByAnchor: Map<number, number[]>): Score {
|
||||
let tp = 0, fp = 0, fn = 0, tn = 0, exact = 0, pairs = 0
|
||||
let noneTp = 0, noneFp = 0, noneFn = 0
|
||||
const falseMerges: Score['false_merges'] = []
|
||||
const missed: Score['missed'] = []
|
||||
for (const a of anchors) {
|
||||
const picks = new Set(picksByAnchor.get(a.rn) ?? [])
|
||||
const truth = new Set(a.candidates.map((c, i) => (c.org === a.anchor_org ? i : -1)).filter((i) => i >= 0))
|
||||
let ok = true
|
||||
a.candidates.forEach((c, i) => {
|
||||
pairs++
|
||||
const p = picks.has(i), t = truth.has(i)
|
||||
if (p && t) tp++
|
||||
else if (p && !t) { fp++; falseMerges.push({ anchor: a.anchor_k, candidate: c.k, sim: c.sim }) }
|
||||
else if (!p && t) { fn++; missed.push({ anchor: a.anchor_k, candidate: c.k, sim: c.sim }) }
|
||||
else tn++
|
||||
if (p !== t) ok = false
|
||||
})
|
||||
if (ok) exact++
|
||||
const predNone = picks.size === 0, truthNone = truth.size === 0
|
||||
if (predNone && truthNone) noneTp++
|
||||
else if (predNone && !truthNone) noneFp++
|
||||
else if (!predNone && truthNone) noneFn++
|
||||
}
|
||||
const r4 = (x: number) => Math.round(x * 10000) / 10000
|
||||
return {
|
||||
anchors: anchors.length,
|
||||
pairs,
|
||||
pair_precision: r4(tp / Math.max(1, tp + fp)),
|
||||
pair_recall: r4(tp / Math.max(1, tp + fn)),
|
||||
pair_tnr: r4(tn / Math.max(1, tn + fp)),
|
||||
anchor_exact: r4(exact / anchors.length),
|
||||
none_precision: r4(noneTp / Math.max(1, noneTp + noneFp)),
|
||||
none_recall: r4(noneTp / Math.max(1, noneTp + noneFn)),
|
||||
false_merges: falseMerges.slice(0, 40),
|
||||
missed: missed.slice(0, 40),
|
||||
}
|
||||
}
|
||||
|
||||
async function main() {
|
||||
const truthNone = anchors.filter((a) => !a.candidates.some((c) => c.org === a.anchor_org)).length
|
||||
console.log(`anchors ${anchors.length}, pairs ${anchors.reduce((s, a) => s + a.candidates.length, 0)}, anchors with no true match ${truthNone}`)
|
||||
const rules = new Map<number, number[]>(anchors.map((a) => [a.rn, rulesSelect(a)]))
|
||||
const report: Record<string, unknown> = { gold: goldPath, rules_v0: score(rules) }
|
||||
if (runLlm) {
|
||||
const picks = new Map<number, number[]>()
|
||||
const usages: unknown[] = []
|
||||
let model = ''
|
||||
let i = 0
|
||||
for (const a of anchors) {
|
||||
const r = await modelSelect(a)
|
||||
picks.set(a.rn, r.picks)
|
||||
usages.push(r.usage)
|
||||
model = r.model
|
||||
if (++i % 25 === 0) console.log(`model: ${i}/${anchors.length}`)
|
||||
}
|
||||
report.model_zero_shot = { model, usages, ...score(picks) }
|
||||
}
|
||||
writeFileSync(outPath, JSON.stringify(report, null, 2))
|
||||
const line = (name: string, s: Score) =>
|
||||
`${name.padEnd(16)} pair P ${s.pair_precision} R ${s.pair_recall} TNR ${s.pair_tnr} | anchor exact ${s.anchor_exact} | none P ${s.none_precision} R ${s.none_recall} (anchors ${s.anchors}, pairs ${s.pairs})`
|
||||
console.log(line('rules_v0', report.rules_v0 as Score))
|
||||
if (runLlm) console.log(line('model_zero_shot', report.model_zero_shot as Score))
|
||||
console.log(`report: ${outPath}`)
|
||||
}
|
||||
|
||||
main().catch((e) => {
|
||||
console.error(e)
|
||||
process.exit(1)
|
||||
})
|
||||
Reference in New Issue
Block a user