Prototype 06 · The Alembic
Better null than wrong

DestilIA
doesn't generate.
It distills.

A multi-stage intelligence pipeline that reads a merchant's raw catalogue — titles, attributes, legacy copy, product photography, the storefront itself — and distills it into verified, on-brand product pages, category landings and brand pages.

Every stage runs on Datexio's own inference fleet. Reasoning models where deliberation pays, fast non-reasoning models where volume does. Every token is audited before it ships.

REASONING + NON-REASONING MODELS ON-PREM GPU FLEET TOKEN-LEVEL AUDIT SOURCE PROVENANCE SELF-HEALING RETRIES

// example · one corrupted token is enough to withhold the page

Product pages distilled
33K+
Model calls audited per night
20K+
Corrupted pages shipped in validation
0
Self-healing tiers
3
// 02 · The Still

Eight specialised stages, one verdict.

Each stage owns a single responsibility and a tailored stack. Nothing reaches the writer that hasn't been read, weighed and cross-checked; nothing leaves the writer that hasn't been audited.

01

Identity Reader

storefront intelligence

Before a single word is written, DestilIA reads the store itself: homepage and "about us", stripped of navigation, prices and promotions. A language model distills a structured store identity profile — who writes for the store, what it sells, to whom, how it addresses the customer, what it emphasises, the core vocabulary of its trade and whether it operates in a regulated sector. That profile conditions every downstream stage: a slot-racing store and a hardware store never sound alike, and a leaf category called "Magnesium" is understood as a slot-car rim, not a mineral. Designed field by field for the model that consumes it.

// ~4 s per storefront

02

Compliance Gate

regulated-vertical classifier

A sub-second classifier decides, per product, whether the item belongs to a regulated vertical — pharmacy, supplements, cosmetics, baby care — and holds it back from automated copy. For food and supplement brands, a dedicated rule block enforces EU Regulation 1924/2006 on nutrition and health claims across every generated field.

03

Vision Interpreter

reasoning vision model + auditor

When the merchant's text is thin, a vision-language model with explicit reasoning reads the product photo the way a buyer reads a shelf: shape, construction, finish, pattern, colour. A second pass audits its own statements and splits them into safe and risky. Anything that requires counting, comparing, interpreting symbols, reading printed text or describing props is quarantined. Only safe observations reach the writer. The photo is always the lowest-trust source: it may describe how a product looks — never what it is made of, how big it is or what the box contains.

04

Coherence Analyst

cross-source contradiction detection

Merchant data is rarely consistent. DestilIA ranks its sources — brand, title, attributes, category, legacy copy, photo — and cross-checks the hard facts that change what the customer receives: brand, model, dimensions, weight, maximum load. When two merchant sources disagree on one of them, the pipeline refuses to write and reports the conflict instead of picking a side. A product page with the wrong measurement is worse than no page.

05

The Writer

single-pass constrained generation

One call, one fixed contract. The writer returns a structured object, not free text: description, body, refusal reason if any, detected conflicts, and a feature list in which every item declares its source — title, description, body, attributes or photo. Code, not the model, assembles the final HTML. Years of experiments collapsed into a single, tightly specified pass that is more accurate and faster than the multi-agent verifier chains it replaced.

// 5–15 s per product page · text-only

06

Taxonomy Cartographer

unsupervised product-type discovery

Category pages are written from what a category actually contains. Titles are cleaned of colour, size, codes and brands, projected into a semantic embedding space, reduced and clustered by density. Language models name each cluster and merge them by purchase intent: two types are the same if a buyer looking for one would accept the other. Discovered types plus real representative products feed a category writer that cites the catalogue's actual ranges. Parent categories are written from their children, with a different editorial role: orient, don't sell.

// ~145 categories in ~20 min · 2 to 30,000+ products per category

07

Brand Voice

brand pages · structural fingerprint

Brand pages describe what this store carries from a brand, not an encyclopaedia entry. Product types are discovered per brand, a stratified sample of titles covers the full range, and the store's legacy text is mined for facts — never for phrasing. A deterministic structural fingerprint, derived from brand and store, assigns each page its own opening, emphasis and layout. The same brand never reads the same in two stores, and search engines never see duplicated content across the network.

// 35 brand pages in under 6 min

08

Integrity Guard

Flagship · token-level audit

The last stage, and the reason DestilIA can publish without a human in the loop. Every response from every model is audited before any other stage sees it. The Guard does not ask where a defect came from: it detects the corruption, whatever its origin, and decides what may leave the building.

▸ Token-level confidence telemetry. The inference fleet exposes the model's own certainty at every generated token. A corrupted token has a signature: confidence collapses in the middle of a word or a number. The Guard reads that signature in real time and catches what no spell-checker or regex would ever see — a colour temperature that loses a digit, a technical standard with a broken number, a brand name that drifts by one syllable.
▸ Cross-script integrity. Characters from writing systems that were not in the input are treated as contamination — including look-alike letters visually identical to Latin ones and invisible to a human reviewer.
▸ Self-healing. A failed response is regenerated through three escalating recovery strategies. Most defects are healed transparently, in seconds.
▸ Quarantine. Whatever cannot be healed never reaches the storefront. Product pages are routed to a review queue; category and brand pages are withheld. Better null than wrong.
// three possible verdicts
Released
clean on every lens
Healed
regenerated · re-audited
Quarantined
review queue · withheld
// guard log · replay run
INTEGRITY GUARD · ON-PREM
> 108 product pages · hardware catalogue · replay run
> 20 flagged at token level · 0 false positives
> corruption caught: "Colour temperature: 640?K"  → quarantined
> corruption caught: "Standard: DIN 309?"          → quarantined
> corruption caught: look-alike glyph in brand name → quarantined
> 88 pages released · 0 corrupted pages shipped

// example output · replay of a real validation run · validated on 5 stores across different verticals · ~0.1% of a standard catalogue routed to human review

// 03 · The Thesis

Not a wrapper around an API. Three things that are hard to copy.

Owned infrastructure

The inference runs in-house.

A heterogeneous on-prem GPU fleet, a model gateway, work queues and routing by job type. Each task gets the model it needs: reasoning where deliberation pays, speed where volume does. Unit cost, latency and data residency are engineering decisions, not a vendor's price list.

Accumulated know-how

Criteria only production teaches.

Which source wins when two disagree. What cannot be asserted from a photo. Which regulation applies to which vertical. How to keep a thousand stores from converging on the same text. None of this is in a model; it is encoded in the pipeline.

Integrity over volume

Unverified output does not ship.

Every model call is audited at token level before any other stage sees it. Refusals are reasoned, conflicts are reported, defects are healed or quarantined. 0 corrupted pages shipped across validation runs on 5 stores in different verticals.

// 04 · The Distillery

Engineering depth, visible at every layer.

⎔

Heterogeneous on-prem GPU fleet

Dedicated accelerators run different model families side by side: reasoning models where deliberation pays, fast non-reasoning models where throughput does. No third-party inference, no data leaving the building.

⇉

Model gateway, dual-lane scheduling

A single API surface routes each job to the right model and accelerator. Vision-heavy and language-heavy work flow through separate lanes, so a long photo-analysis batch never blocks category or brand work.

◈

Provenance by design

Every generated attribute declares its origin. Every refusal carries its reason. Every run leaves a review report next to the importable file.

⟲

Resumable, cache-first jobs

Any job can be paused, cancelled or resumed without repeating finished work. A multi-hour catalogue run survives restarts transparently.

≡

Deterministic where it matters

Product-type discovery and structural assignments are reproducible: the same catalogue yields the same types and the same page structure. Controlled variance is reserved for wording only.

⛬

Built for the platform

Native to OpenTiendas: ingests the platform's own exports, returns ready-to-import files, respects the store's category tree and never touches languages or fields it wasn't asked to.

/// /// /// /// ///

Raw catalogues in.
Distilled truth out.
Every store has a voice. DestilIA writes in it — and audits every word.

The team runs its own inference fleet and ships production-grade AI pipelines daily: from raw merchant data to verified, publish-ready content.

// next: multilingual distillation across storefront languages, audited with the same token-level guard.