Introducing THEMIS: Europe’s Enterprise-Ready System One Model for AI Governance

Colchix is building THEMIS, a sovereign System One model for real-time data protection and AI governance in Europe. It begins with one demanding enterprise task: finding sensitive information, understanding whether it requires protection and returning an allow, mask, block or escalate decision before data reaches an external AI system.
Large language models changed how people interact with software. System One models are beginning to change how software itself makes decisions.
Jev demonstrated the potential of models designed for fast, typed and probabilistic judgments instead of open-ended text generation. The emerging open-source ecosystem has shown that similar decision models can also run locally.
But European enterprises still face a gap.
The models that decide quickly do not necessarily locate the exact information that must be protected. The systems that locate sensitive values may not judge them in context. General-purpose LLMs can perform both tasks, but may be too slow or operationally expensive for an invisible protection layer that must evaluate every prompt and attachment in real time.
And model capability alone does not provide European data residency, zero data retention, private deployment, policy enforcement or audit evidence.
THEMIS is being developed to bring those requirements together.
Starting with PII. Building Europe’s decision layer for enterprise AI.
THEMIS is currently under development. This article describes its intended product direction and the measurable acceptance criteria Colchix has established before release. It does not present projected performance as achieved benchmark results.
What is THEMIS?
THEMIS is Colchix’s sovereign, enterprise-ready System One model for real-time data protection and AI governance.
Its first use case is runtime sensitive-data protection. Given a prompt, message, document or application state, THEMIS is being designed to:
- identify the exact boundaries of personal, confidential and regulated information;
- determine whether each value requires protection in its context;
- return a bounded allow, mask, block or escalate decision;
- expose useful confidence so uncertain cases can be handled safely;
- operate as part of a European-controlled governance and audit layer.
This is different from asking a chat model to generate a compliance assessment. THEMIS is intended to return machine-readable decisions that software can act on before information leaves an approved environment.
The initial model is focused on sensitive data because runtime masking exposes the hardest combination of requirements: precise location, contextual judgment, calibrated uncertainty, whole-file reach and low latency.
The longer-term direction is broader: a general European decision model and specialised System One models for regulated enterprise workflows.
Why existing approaches leave a gap
Protecting an enterprise AI interaction requires more than recognising that sensitive information exists somewhere in the text.
A production system must answer four questions:
- What exactly is sensitive? The complete value must be located, not merely detected.
- Does it require protection here? Context determines whether a value is sensitive or a harmless look-alike.
- Can the entire input be evaluated? Contracts, exports and code files cannot be silently truncated.
- Can the decision happen before transmission? A protection layer must operate without destroying the user experience.
Different technologies solve different parts of this problem.
Pattern matching is fast but narrow
Regular expressions are deterministic, inexpensive and excellent for known formats. They can identify many email addresses, phone numbers, identification numbers and financial patterns.
They cannot reliably understand a project codename, a medical condition described in ordinary language, a person referred to indirectly or a company-specific sensitive term that does not follow a fixed format.
Entity extraction models locate values but may lack judgment
PII and named-entity models can return spans inside text. They are useful when the categories are known and the values resemble their training data.
But locating a possible entity is not the same as deciding whether masking it is necessary or whether masking would break the request. Boundary errors also matter: replacing only part of an address, health condition or identifier still leaks the original information.
Current System One models decide but do not necessarily locate
Decision models such as Jev are designed to return typed judgments and probabilities. That is valuable for classification, routing and policy checks.
However, deciding that a message contains sensitive information does not reveal the exact characters that a masking system must replace. Open decision models may also operate with limited context windows or require additional serving and governance infrastructure.
General-purpose LLMs can reach more text, but slowly
A general LLM can often find values, judge them in context and return structured output. It can also process categories that were not anticipated by a fixed pattern library.
The trade-off is latency, cost and operational dependence on a generative system. A runtime protection layer must evaluate interactions continuously, often without users noticing that another model call has occurred.
What Colchix measured
To define the release bar for THEMIS, Colchix built an internal synthetic benchmark for runtime masking.
The current test set contains:
- 98 workplace-style passages;
- 241 annotated sensitive values;
- 137 harmless look-alikes;
- examples in English, German, French and Italian;
- short messages and long inputs, including approximately 16,000-token documents and exports.
No customer text is included in the benchmark.
The benchmark measures whether a system:
- finds every relevant value;
- returns the exact boundary required for masking;
- judges genuine sensitive values correctly;
- leaves harmless look-alikes untouched;
- reads the complete input rather than silently truncating it;
- completes the task within a usable latency;
- keeps the text inside the required jurisdiction.
The results below are measurements of currently available approaches. They are not THEMIS results.
Current benchmark: every available option trades something away
On the short-message test, the measured systems protected the following share of 100 sensitive values completely:
| Approach | Sensitive values fully protected | Median time per short message | Main limitation |
|---|---|---|---|
| Regex patterns | 59 of 100 | Under 1 ms | Finds fixed formats only |
| GLiNER2-PII | 70 of 100 | 313 ms | Locates known entity classes but does not provide the complete contextual decision layer |
| GLiNER2-PII followed by Jev | 69 of 100 | 697 ms | Chaining two models compounds missed locations and boundary errors |
| Gemini 2.5 Flash, one call | 77 of 100 | 769 ms | Most protective measured option, but also the slowest short-message option in the comparison |
Jev and Laya were also evaluated as decision models. Both can judge text, but neither independently returns the exact spans required to replace every sensitive value. In the tested setup, Jev took a median 384 ms to read and judge a short message; Laya took approximately 480 ms and was shown only the beginning of longer inputs according to the tested configuration.
Long inputs expose another trade-off.
Regex processed the tested long input quickly but protected only fixed-format values. GLiNER2-PII protected 84 of 100 annotated values across the approximately 16,000-token attachment test, but the measured workflow took 119 seconds on the local CPU configuration. Gemini completed one long prose document in approximately 17 seconds but returned no result on the customer-style export after 40 minutes in the tested configuration.
These figures should not be treated as universal vendor benchmarks. They describe a specific synthetic test set, model versions, prompts, local hardware and service configurations. Performance may change with different data, infrastructure, settings and provider updates.
The important finding is architectural:
The fastest options do not provide sufficient semantic protection, while the most protective measured option is too slow for an invisible real-time layer.
Why exact boundaries matter
A value is protected only when the entire sensitive span is replaced.
If a system masks only part of “Keizersgracht 123, Amsterdam,” the remaining text may still reveal the address. The same applies to a medical condition, an internal project name or an indirect description such as “the woman who joined from the Rotterdam team.”
In the measured two-model chain, the largest loss occurred before the final judgment:
- 100 sensitive values entered the workflow;
- 90 were located at all;
- 70 had an exact boundary;
- 69 were finally judged sensitive and protected.
That means 31 of every 100 values still leaked, primarily because values were never found or were only partially located.
THEMIS is being designed to train location, boundary precision, judgment and confidence together rather than treating them as disconnected stages.
The THEMIS release bar
Colchix will not describe projected numbers as achieved results. Instead, we are publishing the minimum engineering acceptance criteria THEMIS must clear before its first production release.
| Acceptance criterion | THEMIS development target |
|---|---|
| Sensitive values fully protected in short messages | At least 90 of 100 |
| Sensitive values fully protected in long attachments | At least 90 of 100 |
| Exact-boundary accuracy | At least 92% |
| Harmless look-alikes left untouched | At least 95% |
| Median processing time for a short message | Under 100 ms |
| Long-input behaviour | Reads the complete supported input, returns an explicit error or escalation, and never cuts silently |
| Data location | European-controlled processing |
| Retention | Zero retention for protected content by design |
These are development targets, not benchmark results. THEMIS has not yet been validated against this acceptance suite.
We will publish measured performance only after the model has been trained, frozen and evaluated against a reproducible test protocol. Where possible, results will include model and dataset versions, execution environment, confidence intervals and known limitations.
Why THEMIS is part of Colchix
THEMIS is not being built as an isolated model demo.
It is intended to become the decision engine inside the Colchix AI governance platform:
- ARGUS discovers which AI tools, accounts and embedded capabilities are being used across the organisation.
- GOLDEN FLEECE intercepts sensitive data before it reaches an external AI system and applies reversible tokenisation or masking.
- THEMIS evaluates what requires protection and returns bounded runtime decisions.
- ATHENA records policies, decisions, confidence, model versions and audit evidence on European infrastructure.
This combination connects model intelligence to an enforceable workflow:
- observe the AI interaction;
- identify and evaluate sensitive information;
- allow, mask, block or escalate;
- send only the permitted content onward;
- restore authorised values locally where required;
- preserve evidence of what happened and why.
The model does not replace deterministic policy. THEMIS supplies bounded judgments; Colchix applies organisational rules, thresholds, access controls and human-review paths around them.
Built for European control
European readiness is not a badge added after training. It affects how the system is designed and operated.
The intended deployment paths for THEMIS include:
- managed serving on European-controlled infrastructure;
- single-tenant deployments;
- private VPC deployment;
- on-premises deployment for organisations with stricter requirements;
- locally controlled masking and token vaults;
- explicit retention and deletion rules;
- pinned model versions and auditable updates.
These deployment paths are product direction and are not all generally available today.
The goal is to let European enterprises use global AI while keeping control of sensitive information, enforcement and evidence within their chosen boundary.
From PII protection to a European decision layer
PII protection is the first proving ground, not the final category.
The THEMIS roadmap is designed to progress in evidence-based stages:
Phase 1: sensitive-data protection
THEMIS operates inside Colchix to locate sensitive information and support runtime masking and policy decisions.
Phase 2: specialised enterprise workflows
Selected design partners define another bounded decision workflow using their own taxonomy and labelled examples. This tests whether the same model approach can generalise beyond PII.
Phase 3: private customer models
Through the planned THEMIS Foundry, a customer’s decision schema and approved data can be used to create a specialised model served privately in Europe.
Phase 4: a broader System One platform
Once repeated customer deployments demonstrate a reliable pattern, Colchix intends to offer a broader European platform for creating and operating specialised decision models.
Each phase depends on proving the one before it. THEMIS Foundry and self-service model creation are roadmap direction, not currently available products.
What belongs to Colchix
THEMIS is being developed on permissively licensed open foundations, with attribution to their respective maintainers.
Colchix is not claiming to have created a new foundation model from scratch. The intended proprietary advantage lies in:
- the labelled enterprise data and category system developed for masking;
- the objective that trains location, decision quality and calibration together;
- the evaluation and acceptance framework;
- European serving and deployment controls;
- integration with Colchix discovery, protection, policy and audit capabilities;
- the additional customer-specific models and workflows built over time.
Open foundations make the starting point accessible. The data, training objective, governance system and production workflow determine whether the result solves an enterprise problem.
Join the THEMIS waitlist
THEMIS is currently under development, and Colchix is selecting European design partners for early technical validation.
The waitlist is intended for organisations that want:
- early access to technical updates and benchmark results;
- an opportunity to evaluate THEMIS on an approved enterprise workflow;
- private European deployment options;
- participation in the design-partner programme;
- future access to specialised System One models.
Join the THEMIS waitlist for technical updates, early access and design-partner opportunities.
Frequently asked questions
Is THEMIS available today?
No. THEMIS is currently under development. Colchix has defined the benchmark, product direction and acceptance criteria, but has not yet published achieved model results or general availability.
Are the THEMIS performance numbers measured?
No. The figures listed in the THEMIS acceptance table are development targets, not measured performance. Benchmark figures attributed to Regex, GLiNER2-PII, Jev, Laya and Gemini describe Colchix’s current internal synthetic evaluation of those existing approaches.
Is THEMIS a fork of Jev?
No. Jev is a proprietary model owned by TypeSafe AI. THEMIS is being developed independently by Colchix on permissively licensed open foundations, with its own labelled data, objectives, evaluation criteria and enterprise integration.
Is THEMIS an LLM?
THEMIS is being designed as a System One decision model rather than a conventional chat or text-generation model. Its purpose is to return bounded, machine-readable judgments and exact sensitive-data locations, not open-ended prose.
Where will THEMIS run?
THEMIS is being designed for European-controlled deployment, with planned managed EU, single-tenant, private VPC and on-premises paths. These options are product direction and are not all generally available today.
What is THEMIS Foundry?
THEMIS Foundry is the planned future capability through which an enterprise could turn its decision schema and approved labelled data into a specialised model served privately in Europe. It is part of the roadmap and is not currently available.
How can an organisation participate?
European enterprises, regulated organisations and technical design partners can join the THEMIS waitlist to receive development updates and discuss early validation opportunities with Colchix.
Methodology and disclosure
The current Colchix masking benchmark uses a frozen synthetic dataset containing 98 workplace-style passages, 241 annotated sensitive values and 137 harmless look-alikes across English, German, French and Italian. The test contains no customer text.
Each existing tool was evaluated on the same underlying examples, subject to the input and output formats supported by that tool. Local timing measurements were taken on a laptop CPU; hosted-model timing includes the measured API request. Long-input results depend on the tested configuration and should not be generalised beyond it.
The dataset was authored internally and currently has one additional blind annotation produced by a model rather than a second human annotator. This is a limitation. Colchix intends to expand the dataset, add independent human review and publish a fuller reproducibility package before making definitive comparative claims.
All THEMIS figures in this article are development targets. No THEMIS performance result is claimed.
Jev and TypeSafe AI belong to their respective owner. Gemini belongs to Google. GLiNER, Laya and other open-source components belong to their respective maintainers. Colchix is not affiliated with or endorsed by these organisations or projects.
Sources
- TypeSafe AI: Introducing System One Models and Jev
- TypeSafe AI technical documentation
- Laya open-source implementation
- Arbiter: local Jev-compatible decision-model serving
- Colchix internal synthetic masking benchmark, version current as of 23 September 2026. A fuller methodology and reproducibility package is planned.
Related: What Is Jev? System One Models vs LLMs Explained
Related: Can European Enterprises Use Jev? GDPR, Data Residency and EU Readiness Explained
Related: Is There a European Alternative to Jev? System One AI, Sovereignty and THEMIS