Latest Decision ·
· · CRF
Continuous Study · Research Day 001

ALEX Research Lab

A live research platform for studying AI governance under uncertainty.

Start Research Session →

Anonymous · No account required · 3-question feedback after each response

Governance Decisions
total since Day 1
Bridge Latency
milliseconds (p50)
Today
decisions so far
Study Status
Active
Human Agreement Rate:
Origin

Why we built this

During thousands of financial conversations, we noticed something uncomfortable. Large language models often sounded equally confident whether the underlying evidence was strong or weak. The fluency was consistent. The epistemic state it reflected was not.

We tried improving the prompts. We tried adding disclaimers. We tried temperature tuning. None of it addressed the core problem: the model was not deciding whether to respond — it was always responding, calibrating only the wording, never the participation itself.

We realized the real problem wasn't generating better answers. It was deciding whether an answer should be generated at all. That observation became the ALEX Governance Runtime.

This research lab exists to test whether that idea works — not in theory, but in the real world, with real users, under real market conditions. Everything here is a measurement, not a claim.

We are not claiming success. We are measuring it.
— The guiding constraint of this research program
Live Decision Stream
Live
Anonymous · No user data · No prompt content
Loading decisions…
The Core Distinction

Model confidence ≠ governance decision

Traditional AI response flow
User query arrives
↓   LLM evaluates
↓   Generates confident-sounding output
Confidence is fluency, not epistemic state
ALEX governance flow
User query arrives
↓   Governance Runtime evaluates signals
↓   Decides: SPEAK / CONSTRAIN / SILENT
LLM only responds when participation is earned

A model that produces fluent, confident text is not the same as a model that should produce that text. ALEX separates the question of participation from the question of generation. The governance runtime runs before the LLM — not as a filter on the output, but as a gate on the input.

Architecture

Five-stage governance pipeline

Before ALEX generates a single token, an independent runtime evaluates the cognitive context and decides whether — and how — to respond.

0
🧱
Posture
HONESTY-BOOT
v1.5
1
🧠
Cognition
ALEX CORE
v2
2
⚖️
Governance
ALEX CORE
v2.6
3
🔀
Deliberation
Intervention
v3.2
4
👁️
Sovereign
ALEX CORE
v3.4
Runtime
V5.0
Policy
v2.6
Thresholds
t1.0
Active since
18 Jul 2026
Full decision path
User Request
signals
Governance Runtime
CRF · IR · CS · OB
decision
Policy Engine
SPEAK · SPEAK_WITH_CONSTRAINTS · REMAIN_SILENT
if SPEAK
LLM
Claude / GPT / any model
governed
Governed Response
+ audit trace_id
"Trustworthy AI requires governance over participation, not only generation."
— The guiding thesis of ALEX Research Lab
Future Vision

From financial AI to governance infrastructure

Finance is the proving ground. The architecture is designed for any regulated AI — healthcare, law, insurance, public administration, enterprise decision systems.

Today
Financial AI Governance
First production research deployment. Public observatory live. Early-stage data collection underway.
Live governance decisions
Public Research Lab
Standalone Governance API
Zenodo pre-registration
Near-term
Governance for Any Regulated AI
Same runtime, applied first to insurance, then to adjacent regulated domains — legal, public sector — with configurable policy profiles per institution.
Domain policy profiles
Enterprise console
Human escalation path
EU AI Act documentation
Future
Independent Governance Infrastructure
Model-agnostic layer above any LLM — Claude, GPT, Gemini, Mistral, internal models. Priced per decision. Institutional scale.
Model-agnostic runtime
Continuous calibration
Peer-reviewed methodology
Open governance standard
Research Timeline

The strategy, not just the roadmap

2025
ALEX Finance — first deployment
Companion AI for financial conversations. First real-world test of the cognitive framework. 25 years of domain expertise encoded.
✓ Delivered
2026 H1
ALEX V5 Architecture — Completed
Five-stage governance pipeline designed and built. First system to treat participation as a computational decision, independent of generation.
✓ Delivered
18 Jul 2026
Production Bridge Deployed
V5 runtime bridge activated in production. First live governance decision recorded. Data collection begins.
✓ Delivered
2026 Q3
Public Observatory + Governance API
Research Lab live. All governance decisions visible in real time. Standalone API open to pilot partners. Study #4 underway.
● Live now
2027
Peer-Review Submission — 2027 AI Governance
Target submission: 2027 workshop on AI governance or trustworthy AI (venue TBD). Requires 10,000 labeled decisions, human calibration feedback, and publication-grade statistical analysis.
In preparation
Future
Enterprise Governance Platform
Policy profiles per institution and jurisdiction. Model-agnostic routing. Continuous calibration. The governance layer becomes infrastructure — not an application, but a standard.
Vision
Live Experiments

Current Research

Study #4 · Human Agreement Study
Does governance improve perceived calibration?
Do users agree with AI silence decisions?
Primary endpoint: user-rated appropriateness of each governance decision
Running
Started
18 Jul 2026
Expected End
31 Oct 2026
Current Sample
Target
10,000
Loading…
Results will be published publicly. Aggregate data visible in real time at /runtime
Measurement Design

What is measured — and what is not

What is measured
User-rated appropriateness — primary endpoint. Did the governance decision match human judgment?
Output mode distribution — SPEAK / SPEAK_WITH_CONSTRAINTS / REMAIN_SILENT rates over time.
Bridge latency — governance overhead at p50 / p95 / p99 in milliseconds.
Signal values — CRF, IR, CS, OB per decision, aggregated anonymously.
What is not measured
Factual accuracy — the governance runtime does not verify the correctness of responses.
Long-term outcomes — downstream effects of governance decisions are outside this study's scope.
Prompt content — no user query content is stored, analysed, or published.
Individual behaviour — all data is aggregated. No user-level tracking or profiling.
Research Contributors

Building this dataset together

Governance decisions recorded
Generated today
ms
Average bridge latency

Thank you. Every conversation helps improve the safety and transparency of AI-assisted decision making. Your anonymous governance feedback contributes to aggregated research statistics. Individual conversations are not published.

Research → Infrastructure

From evidence to infrastructure

The Research Lab is not a standalone project — it is the evidentiary foundation for a stack of products, each building on the findings of the one before.

Research Lab
Evidence
Does the governance runtime align with human judgment? Pre-registered methodology. Empirical validation in progress.
Observatory
Transparency
All decisions public and aggregated in real time. Reproducibility through versioned cohorts and public API.
Governance API
Product
Place the runtime in front of any LLM. Pilot partners now. Growing toward commercial availability.
Enterprise Platform
Infrastructure
Policy profiles per institution. Model-agnostic. Article 9 evidence packages. The governance layer as a standard.
Request Governance API Access →
Open Data

Research Resources

📄
Technical Research Memorandum
TACA protocol, pre-registered methodology, Protocol Lock v1
Zenodo DOI 10.5281/zenodo.20800461
📡
API Documentation
Public governance API — summary, stats, recent decisions, version cohorts
● Live · No auth required
🔭
Live Observatory
Real-time aggregate statistics, latency percentiles, signal averages
● Updating every 2 min
📊
Governance Schema
Anonymous decision records — CRF, IR, CS, OB, output_mode, latency_ms
● Live JSON
🗄️
Dataset Status
Aggregate dataset for download, updated monthly, anonymized and validated
Planned · Q4 2026
⚙️
Standalone Governance API
Place ALEX in front of your own LLM. Model-agnostic. Free tier available.
◐ Pilot · Design partners
⚠️Known Limitations

This section is not a disclaimer. It is the most important part of this page. Honest disclosure of limitations is more valuable to the field than any claim of achievement.

L1

CRF is a proxy metric. Computed from VIX and regime confidence — not a direct measure of epistemic uncertainty.

L2

Output modes are under active evaluation. Thresholds were set a priori and are being validated against user feedback for the first time. This study is that validation.

L3

The governance runtime does not guarantee correctness. SPEAK means the system believes a response is appropriate — not that it is factually accurate.

L4

User labels carry self-selection bias. Participants who provide feedback may differ systematically from those who do not.

L5

Domain is financial Q&A only. Governance behaviour may differ substantially in other domains. Generalization requires additional studies.