Hypotheses on record
Four claims, stated and time-stamped before collection began. What would count as failure is written down too.
AI claims your auditors can verify — pre-registered, in production.
Most AI vendors show you a benchmark; mnemo shows you a method. The Memento-Skills learning loop is being replicated under real multi-tenant production traffic, with four hypotheses locked before the first byte of data flowed. When the study ends, the full replication package opens — so the results you buy on are results anyone can check.
Method / pre-registration
Pre-registration removes the vendor’s favourite trick: deciding what success means after seeing the results. Here the questions, metrics and analysis were fixed first — publicly.
Four claims, stated and time-stamped before collection began. What would count as failure is written down too.
The study runs on live multi-tenant workloads — the messy conditions your deployment will actually face.
Telemetry is anonymised at the source. Evidence accumulates; personal data does not.
Eight LLM providers under the same protocol. Results describe the method, not one vendor’s good week.
Fifteen personas learn and change during the study — the growth curve is the measurement, not an anecdote.
Protocol, anonymised data and analysis code open at study end. Reproduce it, or hire someone who will.
The protocol
Hypotheses, metrics and window locked and time-stamped.
28 days of live production traffic, anonymised at source.
Only the pre-registered analysis — no fishing expeditions.
Replication package published: protocol, data, code.
What this buys you
Enterprise AI purchases fail on unverifiable promises. A pre-registered study is a promise with a receipt.
No moving goalposts.
Success criteria were fixed in advance — results cannot be reframed after the fact.
Procurement-ready.
A locked protocol and open package give due-diligence teams something to actually diligence.
Negative results included.
Hypotheses that fail are published with the ones that hold. That is what makes the ones that hold worth something.
Independently checkable.
Your data scientists can rerun the analysis themselves — no trust in the author required.
Benchmarks measure a model on a Tuesday. This measures a method over 28 days of real work.
Live traffic, real tenants, genuine noise. The conditions are the point.
The protocol allows the study to fail publicly. Marketing departments do not sign up for that.
The study is live and the protocol is public. Briefings on method and interim status go through the main page.