The first form extracts perfectly. Then the address moves, the label changes, or the next document calls it something else. We add another rule. This series explores how to retain what those exceptions teach us: the structure we want, the evidence that supports it, and reusable patterns for finding it again.
Structured Extraction series
Part 1: why extraction rules become fragile, and how evidence, field neighbourhoods and semantic templates can help them generalise.
Next: teaching through the workbench, comparing inference layers and measuring the effect of reviewed examples.
Source: document-ledger. The running example uses Morph and Avalonia; the approach extends the local template-learning ideas in StyloExtract.
Suppose I need to import property applications. The first form has an address box in the upper-right corner, a requested amount beside “Loan amount”, and an applicant name near the top. I can make that example work with a few rules: read this rectangle, find that label, parse the number that follows it.
Those are sensible rules for a stable form. The trouble starts when the documents vary.
A revised application moves the address below the applicant details. A valuation report calls it “Subject property”. A letter gives the same address inside a paragraph. Another form has three address sections, each containing text that passes the postal-address parser.
Now the rectangle selects the wrong area. The exact label is absent. The flattened text has lost the relationship between a heading and its value. “Find an address” succeeds, but returns the address where correspondence should be sent rather than the property being valued.
A familiar response is to add exceptions: another label, another rectangle, another document-specific branch. Each patch captures something useful, but that knowledge is spread across selectors, parsing code and assumptions about particular files. Changing a broad rule can fix one form and damage another.
The expensive failure is a plausible value assigned to the wrong field. A missing address is easy to flag. A valid correspondence address silently imported as the property address can pass basic validation.
This is the problem I want to explore: how do we keep useful extraction knowledge without making every new document another special case?
An LLM can help when exact labels and fixed positions fail. Given the text and enough context, it may recognise that “Subject property” supplies the address we want. A model that can see the page can also use visual grouping that a flattened text stream has lost.
That helps propose a mapping. We still need to establish which source supports it, whether a competing address was overlooked, and whether several mentions describe the same property. Producing JSON with an address member establishes the output shape; the value's role still needs evidence.
There is also the next import. If a person corrects the result and we save only the corrected string, we have repaired one record. The useful discovery — “this address belongs to the security-property section” — has gone missing. A later model call or rule run has no retained explanation to reuse unless we explicitly provide one.
So the proposed solution needs two kinds of memory: the facts recovered from documents, and the patterns that helped recover them. Rules and models can both contribute to that memory.
That is the idea behind a live semantic template.
Reviewed corrections generate learning proposals. In this design, a shared template changes only after a proposed revision has been reviewed, qualified on other documents and explicitly adopted. “Live” describes how evidence can inform its evolution; production templates are versioned and do not silently rewrite themselves.
The template defines the structure we want: fields, types, required values and relationships. Alongside that structure, it can retain several ways a field is expressed: labels, neighbouring fields, section membership and relative positions.
The ledger records what we observed: the source document, the text or region, the proposed field role, the parsed value and any review decision. Its positions identify evidence we can inspect. They also supply features from which reusable spatial relationships can be learned.
Together they let a correction carry more information than a replacement value. Confirming a property address can retain the supporting region, the heading that scoped it and the correspondence address that was rejected as the wrong role. A later document can be compared with that reviewed neighbourhood.
The reusable unit can be a field or content area. An unfamiliar document may share a familiar address block even when the rest of its layout is new. Several recurring layouts can become variants of one semantic field, with their applicability inferred from context. We do not need a separate, isolated template for every document fingerprint.
The ledger can record a correction immediately. Promoting a reusable pattern is a separate decision: a useful fact about this document may be too specific to improve extraction elsewhere. The feedback may come from a person, a model comparison or a failed validation check; retain who or what supplied it and the source evidence it concerns.
Across the application, valuation report and letter, the output we need is comparatively stable: an applicant, a property and a requested amount. The input layouts vary much more than that target structure.
Here is a deliberately small conceptual schema. It states what a successful import must recover; extraction patterns will describe where to look for supporting evidence.
schema: property-application
fields:
applicant.name:
type: person-name
required: true
cardinality: one
property.address:
type: postal-address
required: true
cardinality: one
labels: [Property address, Subject property]
components: [unit, building, street, locality, postcode, country]
coalesce:
scope: [entity, semantic-role, effective-period]
preserve: [source-values, evidence, conflicts]
loan.requested_amount:
type: money
required: true
cardinality: one
currency: GBP
The schema gives the extractor a concrete job: find evidence for these roles, parse it into the required types and expose missing or ambiguous fields. Three mentions of one property address can support one resolved field. An absent required value stays missing.

The concrete example: a synthetic application rendered by Morph, with field values linked back to their source regions. The highlighted text supplies evidence for the structured record.
The screenshot makes that separation visible. The record on the right has named fields; the document on the left supplies the evidence. A successful extraction must connect the two, including when the next document changes its wording or layout.
There is already a useful example of this approach in HTML. In StyloExtract, the target roles include Title, MainContent and PrimaryNavigation. The input is arbitrary HTML. The extractor identifies useful structure, learns selectors and reuses them across pages.
Its cached rules avoid volatile class names, sibling indexes and page text. Approximate structural fingerprints retrieve reusable extractors as the content changes.
Documents add another set of clues. A label can sit beside its value. A heading scopes several fields. A table column supplies meaning to every cell below it. A bordered group can separate the property address from the applicant's current address.
That is the relationship worth carrying across to document extraction: identify useful structure, retain how to find it, then reuse it when the content changes. A fingerprint can retrieve a familiar document pattern. A semantic field also needs reusable local patterns, because the same address neighbourhood can occur inside several otherwise different document types.
Return to the rule that found a valid address in the wrong section. A form may contain three perfectly valid candidates:
| Source context | Intended role |
|---|---|
| Applicant details → Current address | applicant.current_address |
| Security property → Address | property.address |
| Correspondence → Send letters to | correspondence.address |
A postal-address recogniser might find all three. Our schema needs to distinguish them. The value's shape tells us that it is an address; its neighbourhood helps tell us whose address it is and what it means.
That neighbourhood includes the label, section heading, nearby fields, relative position and any containing table or group. “Purchase price” and “Property type” are useful neighbours for a property address. “Time at current address” points towards an applicant's address.
Positions become semantic evidence through those relationships. “Below the security-property heading, inside the same group” can survive a page redesign. A rectangle at a fixed page coordinate is much more dependent on that particular layout.
A graph can represent these relationships: blocks, labels, cells and sections are nodes; containment, alignment and adjacency are edges. HTML supplies DOM relationships. PDF geometry lets us infer many of them.
Even after we locate the right role, collecting every match independently can give us the wrong structure. One document says “Property address”. Another says “Subject property”. A third splits the same address into street, town and postcode boxes. Importing each as a new field loses the fact that they may all describe the same property address.
The ledger needs semantic deduplication and coalescing: several pieces of compatible evidence can support one canonical field, while their original observations remain inspectable.
For the same reviewed property and address role, “42 Example St, Edinburgh, EH11AA” and “42 Example Street, Edinburgh, EH1 1AA” may be equivalent after address-aware normalisation. Separate street, town and postcode boxes may supply components of that same value. The resolved field keeps the raw observations and the source region that supplied each component.
The correspondence address remains a distinct role even if its text is identical. A conflicting flat number must also remain visible. Similar wording alone cannot establish shared identity, and a missing component stays missing until compatible evidence supplies it.
This is part of structured extraction: three mentions should not accidentally become three properties. Coalescing needs an inspectable reason and a reversible decision. The core contract and implementation carry the detailed rules for component assembly, scope, conflicts and source lineage.

Five property mentions/components contribute to one resolved address with five source regions. The same text in a correspondence section retains its separate role. Flat 2 and Flat 3 remain visible alternatives. Turning off automatic coalescing exposes all eight observations.
The first few examples can look like one pattern with a growing list of exceptions. To learn something reusable, we need to distinguish changes in presentation from changes in meaning. “Split the field” can describe three different operations.
A single block might read:
Alex Morgan · alex@example.test · 07123 456789
A target schema may ask for a name, email and telephone number. Segmentation proposes spans; parsers help assign their types. Conversely, several spans can assemble into one address. Repeated rows need separate boundaries when each describes a different property or line item.
The same property.address can appear in an application box or a valuation table. These are variants of one semantic field: different labels and geometry, shared type and meaning.
Property and correspondence addresses have distinct roles even when their text is identical. This is a semantic distinction supplied by the schema and reviewed examples. A reviewer correcting the selected role supplies information that clustering similar text cannot establish by itself.
To compare those patterns, we need a representation of the context. In ML terminology, a feature is a measurable description of an example. For a field neighbourhood, useful feature groups include:
| Feature group | Example |
|---|---|
| Text | Label wording, section heading, nearby labels. |
| Geometry | Value-to-label offset, alignment, containment, page-relative bounds. |
| Structure | Table row/column, DOM ancestry, repeated-group membership. |
| Value shape | Multiline address, currency amount, date, email. |
| Document context | Headings and other fields suggesting a document or section type. |
An embedding turns text or an image into a numerical representation that can support similarity search. It can help connect “Subject property” with “Property address” when exact string matching fails. An image embedding may contribute information about boxes, grouping and visual layout.
We still need the explicit features. A text embedding might retrieve both property and correspondence addresses because their neighbourhoods are linguistically similar. Section relationships, competing roles and reviewed negative examples help resolve that ambiguity.
The features need to capture recurring context. Customer names and particular amounts should not become template constants. Their shapes may help parse a field, while its label and neighbours help identify its role.
The first reviewed example gives us a neighbourhood seed: a label, nearby text, positions and a value linked to source evidence. It is one observed example. It contains no computed embedding vector or statistical cluster estimate in the current prototype.
With more reviewed examples, we can look for recurring patterns. A prototype represents a pattern; a centroid is an average position in a chosen numerical feature space. That statistical summary becomes useful only once we have defined and computed the features being averaged.
One field may need several prototypes. If applications put the address below its label and valuation reports put it to the right, averaging those positions can describe a location that works for neither. Separate variants can retain the common semantic role while representing each recurring layout.
The challenge is to avoid recreating the original pile of exceptions. In ML terms, an overly broad pattern can underfit, missing useful distinctions. A template for every individual document can overfit, memorising incidental details. A proposed variant should earn its place by improving extraction on documents held apart from those that suggested it.
Document type helps choose a variant, and familiar fields help suggest a document type. An unfamiliar form can therefore start from shared field knowledge while its type remains uncertain. The next article will explore how reviewed examples justify extending a pattern or proposing a split.
The overall flow is:
flowchart TD
A[Template: fields, types and relationships] --> D[Match evidence to semantic roles]
B[Input: text, regions and structure] --> C[Candidate spans and neighbourhoods]
C --> D
E[Shared patterns and field variants] --> D
D --> F[Parse values and assemble components]
F --> J[Deduplicate and coalesce compatible observations]
J --> K[Validate resolved fields and expose conflicts]
K --> G[Ledger observations with source evidence]
G --> H[Human and model feedback proposes pattern revisions]
H --> L[Review and qualify on other documents]
L --> M[Explicitly adopt a versioned template revision]
M --> E
Explicit output constraints remain useful throughout: a money parser checks the amount, section exclusions can reject a tempting wrong role, and cardinality checks the field instances after equivalent observations have been coalesced. Ambiguous identity or incompatible values can stay unresolved.
Each improvement now addresses a failure from the opening. A reviewed “Subject property” label adds vocabulary. Correcting correspondence to property teaches a section relationship and retains a negative example. Repeated table layouts can justify a specialised variant. Compatible observations can supply address components without importing duplicate properties.
These are benefits to measure. More templates or a higher similarity score do not themselves establish better extraction. Use documents held apart from the teaching examples, and check correct role-and-value mappings, missed fields, false merges and review workload. Reimporting the same file must not inflate evidence for a pattern.
There is a possible cost benefit too: a qualified pattern may let familiar documents use deterministic extraction, retaining a model's useful discovery as a reusable rule.
That gives us a practical reason to retain a ledger: the next improvement should be traceable to the examples and decisions that justified it.
Consider the competing addresses again. An embedding retriever supplies plausible neighbourhoods. A small decision model compares their roles. A parser checks the address components. A person brings knowledge of the case and can correct the proposed mapping. These components answer different questions and contribute different kinds of evidence.
Their scores are not interchangeable. Embedding similarity describes proximity in a representation; a role classifier may estimate a probability; a parser reports whether a value meets a rule. We cannot average those numbers and call the result confidence. Calibration asks whether probability estimates agree with observed correctness on comparable cases. The model's stated confidence needs that check too. On Calibration of Modern Neural Networks.
Human review belongs inside this loop. Selecting the property address and rejecting the correspondence address supplies a positive example, a negative example and a preference between candidates. People can also disagree or make mistakes, so preserve the reviewer, rationale and evidence. An explicit review decision can govern this record while its proposed generalisation still needs qualification.
The ledger keeps these signals distinguishable. Source checks and role/component constraints govern what can be merged and emitted. Model agreement can support a proposal, but two models reading the same region do not create two independent source confirmations. Unresolved conflicts can return to a person.
A corrected role is a supervised label. Choosing one candidate over another supplies preference feedback. A validation failure is a weaker signal: it shows something is wrong, but may not identify the correct answer. The chosen neighbourhood can become a positive retrieval example; the rejected address a negative. The reviewed role can label or evaluate a candidate ranker. A verified extraction outcome could later reward a policy that chooses which extractor to run or when to request review. One correction can therefore inform several components, each through an appropriate learning signal.
Reinforcement learning from human feedback (RLHF) uses human feedback to form a reward signal for optimising a model's policy. The familiar language-model approach learns rewards from human preferences and then optimises against them. Training language models to follow instructions with human feedback.
The current workbench records feedback and neighbourhood seeds; it does not run that training loop. The immediate learning path is to retain labelled evidence, propose a reusable change and qualify it. A future policy for selecting extractors or requesting review could use reinforcement learning, with success judged by verified extraction quality and review cost. That would be another component whose proposals and outcomes the ledger records.
The next part will compare these contributions through the workbench. The research notes cover candidate embedding and decision models; the common requirement is to preserve role, provenance and negative evidence as feedback moves between components.
The example repository starts with the part this architecture depends on: a proposed field must stay connected to inspectable source evidence. Its Avalonia viewer uses Morph to turn Word and HTML into PDF, then PDFium to render pages and locate text. The source region and the displayed page therefore refer to the same representation.
A person can create fields on a new document, use a local LLM to propose a draft, and correct the values and evidence. The current implementation retains neighbourhood seeds, supports simple label replay and implements the deterministic address-coalescing baseline. Learning generalised variants remains the next stage.
The current neighbourhood seed retains the label, nearby text, relative positions and pending embedding inputs. In the prototype source its type is still named InitialCentroid; that name anticipates later work. No numerical centroid is computed here.

The inspectable seed: a reviewed field retains the label, surrounding text and source positions. This is the evidence from which a reusable pattern can grow.
This actual record from DemoLedger.cs connects a reviewed field and its neighbourhood seed to the source:
public sealed class ReviewedNeighborhood
{
public string Field { get; set; } = "";
public string SourceHash { get; set; } = "";
public string Anchor { get; set; } = "";
public string Section { get; set; } = "";
public InitialCentroid Features { get; set; } = new();
}
MainWindow.axaml.cs captures those seeds and demonstrates basic label replay. LocalTemplateStarter.cs supplies the local LLM proposal adapter. They make the retained evidence concrete; qualification and promotion of learned variants are follow-on work.
The address coalescer now implements a deterministic baseline. It assembles reviewed component blocks before comparing complete values, checks group members for contradictions, and applies cardinality to the resolved fields. One of its first guards is deliberately ordinary code:
private static string? ScopeContradiction(FieldScope a, FieldScope b)
{
if (a.Entity.Length > 0 && b.Entity.Length > 0 && a.Entity != b.Entity) return "Different entities.";
if (a.Role.Length > 0 && b.Role.Length > 0 && a.Role != b.Role) return "Different semantic roles.";
if (a.Period.Length > 0 && b.Period.Length > 0 && a.Period != b.Period) return "Different effective periods.";
return null;
}
Passing that guard only means there is no known scope contradiction. Automatic matching also requires known scope, checked source evidence and compatible parsed components. A missing unit remains unknown. A reviewer can supply an unresolved scope, while known roles and conflicting components remain protected.
This first implementation uses conservative UK-style parsing and explicitly reviewed entity/period context. Inferring those associations and qualifying reusable field variants remain follow-on work.
The starting problem was an extractor that accumulates fragile exceptions. The proposed route forward is to retain what each successful extraction and correction tells us: the semantic role, its source neighbourhood, the pattern that applied and the alternatives that failed.
A template defines the required structure and reusable field patterns. A ledger keeps the observations and decisions that support them. Together they give us a basis for generalising across documents and for checking whether that generalisation actually helps.
The next part follows a correction through the workbench: capture its evidence and feedback, decide whether to extend a pattern or propose a variant, and compare the resulting extraction on documents held apart from teaching. Layer switches will let us examine what embeddings, decisions and human review each add.
The test is whether the next unfamiliar form needs less repair, with the evidence still available to check why.
© 2026 Scott Galloway — Unlicense — All content and source code on this site is free to use, copy, modify, and sell.