Deposit your own documents
Bring into your organization's space a norm, a law or a regulation you hold the rights to, from the app or from the agent you already use.
Stratta does not sell documents. A document (norm, law, regulation, framework) enters your
organization's space because you deposit it yourself, and it stays in your space, readable by your
organization alone. There are two ways to bring it in: from the app, the simpler one, or from
your own agent using the ingest_* tools, so the PDF stays on your machine.
Depositing is included in your plan: it charges no queries. Pro counts 15 norms, Max counts 60, and an extra norm costs CHF 5, bought once. The free plan deposits nothing: it reads the demonstration norm. See Plan and quotas.
The watermark and copyright notice of your copy are kept once with the document and accompany every excerpt served from it: section, citation, assistant answer, deliverable, export.
The rights are your responsibility. With every deposit you declare that you hold the rights to the document: a licence, the publisher's authorization, or a public-domain text (a law, a published decision). SIA standards, Eurocodes and accounting standards are copyrighted, and your organization answers for what it deposits. See the Terms.
Deposit from the app
This is the recommended way: nothing to install, no Python, no agent to connect. You drop the PDF, Stratta reads it on its own server, and a person on the team reviews the result before it joins your corpus.
Open the form
On the Norms page, click Ajouter une norme (Add a norm). The organization owner and admins can do it; members read, they do not add.
Drop the PDF
One file per norm, 200 MB at most, five norms being added at a time. A scanned PDF is accepted: it goes through optical character recognition. The page count is read in your browser, and the dialog tells you whether the text is native or will have to be recognised: a scan costs twice as much and is more likely to get a symbol or a table cell wrong. If you have the published edition of the norm, use it.
Check the fields
The reference (the code), the edition year and the title are prefilled from the file name; correct them if needed. Pick the language (French, German, Italian or English).
Declare the rights
Tick the box confirming you hold the rights to this document (a licence, the publisher's authorization, or a public-domain text). If the document is already in your corpus, the form offers Remplacer (Replace): dossiers that cite it are then marked "to review".
Send
Click to deposit: the document joins the queue.
How long it takes
The window shows the reading time, about two minutes plus six seconds a page (twice that for a scan), so a quarter of an hour for a 117-page document. Stratta's review comes after; an e-mail tells you when the document is ready.
A deposit charges no queries: it is included in the plan. It counts as one norm in the number your plan allows. If you have no room left, delete a document, move to a wider plan, or buy an extra norm (CHF 5, once). See Plan and quotas.
What happens next
The file is processed on Stratta's own server in Germany (EU):
- Reading: optical character recognition if the PDF is scanned.
- Structure: the chapter tree and a summary per section.
- Formulas: formulas and tables are transcribed by a language model.
- Check: a second model checks every formula and table against the image of its page.
- Review: a person at Stratta examines the result before it is added.
A card on the Norms page follows each deposit live: Déposée, Lecture, Structure, Formules, Contrôle, Relecture, Disponible. You can close the page, processing carries on. The time depends on the size of the norm and on the review. When the card shows En relecture (in review), reading is done: the norm is waiting for someone at Stratta to approve it before it joins your corpus.
Withdrawing, hiding or depositing again, and the email
While the deposit is waiting in the queue, Retirer (Withdraw) cancels it. A finished deposit stays on the page for a week; Masquer (Hide) removes it. A deposit that did not go through offers Redéposer (Deposit again): the reference, edition and title are kept, you only drop the PDF. When the norm is in your corpus, or when it is refused or unreadable, the person who deposited it gets an email saying so, with the reasons where relevant.
A deposit that does not go through says why:
| Label | What happened | What to do |
|---|---|---|
| Illisible (unreadable) | The PDF is damaged, password-protected or longer than 2,000 pages. | Deposit an unprotected copy, or split it. |
| Hors forfait (over plan) | Your plan has no room left for one more document. | Delete a document, change plan, or buy an extra norm. |
| Interrompue (interrupted) | Processing failed on our side. Your file is not the cause. | Deposit the same file again. |
What happens to the PDF
The PDF is deleted from Stratta as soon as the document is delivered, refused or withdrawn. The text stays in
your organization's corpus, readable by that organization only. The models are reached through
OpenRouter: every call requires a zero-retention provider (zdr) that does not collect the data
(data_collection: deny), and DeepSeek's own endpoint is excluded. See Security & data.
Or from your agent, PDF on your machine
With the ingest-norm skill, the PDF never leaves your machine: only the result of the ingestion is
written to your workspace. This way needs an agent with a machine and Python, and it runs on your own
subscription's tokens. The rest of this page describes it.
How it works
Ingestion has two passes. A Python pre-pass (ingest-prepass.py, built on PyMuPDF) reads the
PDF deterministically: the table of contents from the bookmarks or the numbered headings, the text
of each section between two headings, and a full-page render of every page that carries a figure.
Then your agent validates the tree, writes a short summary per section, reads formulas and
complex tables visually (automatic extraction mangles both), declares cross-references, and writes
everything to your workspace through the ingest_* tools. ingest_publish scores what actually
made it in and refuses a hollow document.
The procedure is served to the agent as the ingest-norm skill: as an MCP resource
(stratta://skills/ingest-norm) on the remote connector, and inside the npm package on the local
server. You never run the tools by hand.
Prerequisites
- An agent with a machine. The pre-pass reads the PDF on a disk, so this works from Claude Code, Codex, Cursor, VS Code or Gemini CLI, either through the remote connector or the local server. Claude on the web and ChatGPT cannot ingest: nothing runs on a machine of yours there.
- Python 3.10 or newer with PyMuPDF:
python -m pip install --user pymupdf. - The owner or admin role in the organization. Members read, they do not ingest.
- The norm PDF on that machine. A scanned PDF with no text layer needs OCR first (
ocrmypdf); the pre-pass reads text, not pixels.
Ingest a norm
From your agent, in a folder where the PDF is reachable, ask:
Ingest the norm at C:\norms\SIA_263_2013.pdf into my Stratta workspace.The agent loads the ingest-norm skill, downloads the pre-pass script if it is not on disk
(https://stratta.ch/ingest-prepass.py), runs it, then walks through the workflow below. A full
norm takes a while: the agent reads every chapter it publishes.
The ingestion workflow
This is what the skill does, step by step. Understanding the flow helps you review the result.
Check for an existing copy
ingest_status { code }reports whether the norm already exists in your workspace and its coverage. To re-ingest, the old copy is removed first withingest_delete.Run the pre-pass
The script writes a JSON manifest: the tree with a stable id, a human path (
4.2.1,Annexe B.1) and a page range per node, the raw text of each section, and a PNG per page that holds a figure. Annexes are kept: they hold the numeric values. The copy's watermark and copyright leave the page text but not the norm: the manifest keeps them once, indoc.licence.Validate and summarise
The agent checks the tree against the PDF (a heading taken for a table row, a chapter that swallowed the next one), merges trivial paragraphs, and writes a specific summary per section that names the key values and formulas.
Create the document
ingest_create_document { code, year, title, language, totalPages, licence }returns adocumentIdused by every following call.licencerepeatsdoc.licenceverbatim.Insert sections in batches
ingest_create_sectionsbulk-inserts the tree (30 to 50 sections per call) and returns a map from each node id to its stored section id.Enrich each section
ingest_attach_formula(LaTeX, read from the page image),ingest_attach_table(headers + rows) andingest_attach_cross_ref(links to other norms) add the technical content.Upload figures
ingest_upload_figurestores each rendered page (PNG, JPEG or WebP, up to 8 MB) and links it to its section.Normalize cross-references
ingest_normalize_cross_refsscans the text for references to other norms (SIA, SN EN, EN, ISO, DIN) and rebuilds the cross-reference index.Publish
ingest_publishscores the document (text per page, pages under a chapter, empty leaves, formulas and tables) and makes it queryable. Below 30/100 it refuses with the reasons;force: truepublishes anyway once you have confirmed.
See the ingestion tools reference for every parameter.
Quality matters
The accuracy of future citations depends on the ingestion:
- Paths and page ranges must be correct. They are the citation itself.
- Summaries drive how the agent navigates the tree, so they should be specific (mention key formulas or values).
- Content must stay faithful to the source and never be invented.
After ingestion, the Norms page of your workspace shows the coverage of each norm (full,
partial, sparse). Spot-check a few sections by asking your agent questions you already know
the answer to, and confirm the citations land on the right pages.