A multi-stage intelligence pipeline that reads a merchant's raw catalogue — titles, attributes, legacy copy, product photography, the storefront itself — and distills it into verified, on-brand product pages, category landings and brand pages.
Every stage runs on Datexio's own inference fleet. Reasoning models where deliberation pays, fast non-reasoning models where volume does. Every token is audited before it ships.
// example · one corrupted token is enough to withhold the page
Each stage owns a single responsibility and a tailored stack. Nothing reaches the writer that hasn't been read, weighed and cross-checked; nothing leaves the writer that hasn't been audited.
raw catalogue ─▸ 01 identity ─▸ 02 compliance ─┬─▸ 03 vision ─▸ 04 coherence ─▸ 05 writer ─┐ ├─▸ 06 cartographer ────────────────────────┼─▸ 08 guard ─┬─▸ release └─▸ 07 brand voice ─────────────────────────┘ └─▸ quarantine // lanes: 03–05 product pages · 06 category pages · 07 brand pages
Before a single word is written, DestilIA reads the store itself: homepage and "about us", stripped of navigation, prices and promotions. A language model distills a structured store identity profile — who writes for the store, what it sells, to whom, how it addresses the customer, what it emphasises, the core vocabulary of its trade and whether it operates in a regulated sector. That profile conditions every downstream stage: a slot-racing store and a hardware store never sound alike, and a leaf category called "Magnesium" is understood as a slot-car rim, not a mineral. Designed field by field for the model that consumes it.
// ~4 s per storefront
A sub-second classifier decides, per product, whether the item belongs to a regulated vertical — pharmacy, supplements, cosmetics, baby care — and holds it back from automated copy. For food and supplement brands, a dedicated rule block enforces EU Regulation 1924/2006 on nutrition and health claims across every generated field.
When the merchant's text is thin, a vision-language model with explicit reasoning reads the product photo the way a buyer reads a shelf: shape, construction, finish, pattern, colour. A second pass audits its own statements and splits them into safe and risky. Anything that requires counting, comparing, interpreting symbols, reading printed text or describing props is quarantined. Only safe observations reach the writer. The photo is always the lowest-trust source: it may describe how a product looks — never what it is made of, how big it is or what the box contains.
Merchant data is rarely consistent. DestilIA ranks its sources — brand, title, attributes, category, legacy copy, photo — and cross-checks the hard facts that change what the customer receives: brand, model, dimensions, weight, maximum load. When two merchant sources disagree on one of them, the pipeline refuses to write and reports the conflict instead of picking a side. A product page with the wrong measurement is worse than no page.
One call, one fixed contract. The writer returns a structured object, not free text: description, body, refusal reason if any, detected conflicts, and a feature list in which every item declares its source — title, description, body, attributes or photo. Code, not the model, assembles the final HTML. Years of experiments collapsed into a single, tightly specified pass that is more accurate and faster than the multi-agent verifier chains it replaced.
// 5–15 s per product page · text-only
Category pages are written from what a category actually contains. Titles are cleaned of colour, size, codes and brands, projected into a semantic embedding space, reduced and clustered by density. Language models name each cluster and merge them by purchase intent: two types are the same if a buyer looking for one would accept the other. Discovered types plus real representative products feed a category writer that cites the catalogue's actual ranges. Parent categories are written from their children, with a different editorial role: orient, don't sell.
// ~145 categories in ~20 min · 2 to 30,000+ products per category
Brand pages describe what this store carries from a brand, not an encyclopaedia entry. Product types are discovered per brand, a stratified sample of titles covers the full range, and the store's legacy text is mined for facts — never for phrasing. A deterministic structural fingerprint, derived from brand and store, assigns each page its own opening, emphasis and layout. The same brand never reads the same in two stores, and search engines never see duplicated content across the network.
// 35 brand pages in under 6 min
The last stage, and the reason DestilIA can publish without a human in the loop. Every response from every model is audited before any other stage sees it. The Guard does not ask where a defect came from: it detects the corruption, whatever its origin, and decides what may leave the building.
> 108 product pages · hardware catalogue · replay run > 20 flagged at token level · 0 false positives > corruption caught: "Colour temperature: 640?K" → quarantined > corruption caught: "Standard: DIN 309?" → quarantined > corruption caught: look-alike glyph in brand name → quarantined > 88 pages released · 0 corrupted pages shipped
// example output · replay of a real validation run · validated on 5 stores across different verticals · ~0.1% of a standard catalogue routed to human review
A heterogeneous on-prem GPU fleet, a model gateway, work queues and routing by job type. Each task gets the model it needs: reasoning where deliberation pays, speed where volume does. Unit cost, latency and data residency are engineering decisions, not a vendor's price list.
Which source wins when two disagree. What cannot be asserted from a photo. Which regulation applies to which vertical. How to keep a thousand stores from converging on the same text. None of this is in a model; it is encoded in the pipeline.
Every model call is audited at token level before any other stage sees it. Refusals are reasoned, conflicts are reported, defects are healed or quarantined. 0 corrupted pages shipped across validation runs on 5 stores in different verticals.
Dedicated accelerators run different model families side by side: reasoning models where deliberation pays, fast non-reasoning models where throughput does. No third-party inference, no data leaving the building.
A single API surface routes each job to the right model and accelerator. Vision-heavy and language-heavy work flow through separate lanes, so a long photo-analysis batch never blocks category or brand work.
Every generated attribute declares its origin. Every refusal carries its reason. Every run leaves a review report next to the importable file.
Any job can be paused, cancelled or resumed without repeating finished work. A multi-hour catalogue run survives restarts transparently.
Product-type discovery and structural assignments are reproducible: the same catalogue yields the same types and the same page structure. Controlled variance is reserved for wording only.
Native to OpenTiendas: ingests the platform's own exports, returns ready-to-import files, respects the store's category tree and never touches languages or fields it wasn't asked to.
Raw catalogues in.
Distilled truth out.
Every store has a voice. DestilIA writes in it — and audits every word.
The team runs its own inference fleet and ships production-grade AI pipelines daily: from raw merchant data to verified, publish-ready content.
// next: multilingual distillation across storefront languages, audited with the same token-level guard.