Back-office automation for a Swiss services firm.
How we adapted Attera's grounded extraction engine to nine non-CSRD document categories, and what the team will see in a single dashboard, with every figure cited back to its source page.
Manual categorisation. Spreadsheet aggregation. PDFs every month.
A Swiss back-office services firm (small team, recurring monthly document load) handles nine categories of financial documents on behalf of their clients. The original workflow: download PDFs, classify each one by eye, copy the headline figures into a spreadsheet, then reconcile across categories at the end of the cycle.
We took the same grounded extraction engine and pointed it at their inbox. Citation-grounded extraction. No third-party LLM API. Same audit-chain discipline.
This page describes the capability we built. Full attribution (client name, real numbers, before/after metrics) will be added once the firm signs off on public details.
Nine document categories. One classifier. One snapshot.
An auto-classification layer assigns each uploaded PDF to one of nine categories. A category-specific extractor pulls the headline fields: amounts, dates, counterparties, references. A snapshot dashboard aggregates the extracted figures across all documents uploaded into the workspace, with full drill-down to source.
Every figure carries a citation chip back to its source page, the same widget visitors see on the homepage CSRD specimen.
The same chip pattern, on a non-CSRD document.
The widget below shows what an extracted supplier invoice looks like inside the workspace. Same per-line citations as the CSRD specimen on the homepage, just pointed at a different document family.
Extracted from the uploaded PDF · cited per line
Values synthesised for this page. Real client documents stay on EU infrastructure we operate and are never shown publicly without written consent.
Upload → classify → extract → snapshot.
The pipeline mirrors the CSRD flow at every step. The only thing that changes per category is which fields the extractor is told to look for.
Intake
Documents are dropped into the workspace via the upload screen, or forwarded by email when that's wired up. Each upload is hashed and deduplicated; re-uploading the same PDF doesn't double-count.
Classify
The local LLM assigns each document to one of the nine categories. Ambiguous documents are flagged for manual review rather than auto-routed, the same refuse-rather-than-guess discipline as the CSRD pipeline.
Extract
A category-specific extractor pulls the headline fields. Numbers are normalised (Swiss apostrophe formatting, EU decimal comma, multi-currency). Every field carries a citation chip back to the source page.
Snapshot
The snapshot dashboard aggregates extracted figures by category and period. Drill into any cell to open the source PDF at the exact cited page. PII fields can be blurred for screen-share contexts.
The pattern fits any service firm with recurring categorised PDFs.
We didn't build anything CSRD-specific to make this work. Underneath is the same engine: DocTR for text extraction, a local LLM for classification and field extraction, the citation chip for grounding, the append-only audit chain for every approval.
Swap "monthly statement" for "client statement" or "loan schedule" and the same pattern runs. Categories are configurable per workspace; the snapshot dashboard adapts to whatever categories you define.
Who this fits
- Fiduciaries and trustees managing recurring filings
- Multi-family offices reconciling monthly statements
- Accounting firms with bulk supplier invoice processing
- Operations teams swimming in categorical PDFs
- Any back-office workflow where the documents arrive predictably and the figures need to be cited
Want to see this run on three categories from your own inbox?
Zeff and Charles on every demo. Bring sample PDFs across a couple of categories. We'll show you the snapshot in under thirty minutes, on hardware we operate in Belgium.