Your LightDec models and Laya, on your own decisions
Your decision models.
The same question.
Every model reads a state, answer typed questions and return a probability for every option, without writing a word. Give them the same decision and see where they agree, how sure they are, and how fast they answer.
00 / setup
Models load when the lab starts
The first start downloads LightDec_Arthur, LightDec_V2, Enterprise Reflux Laya V2.1 and Laya from Hugging Face (about 2 GB) into the cache volume. Every LightDec (FalconDec) and Arthur model in the local models folder is found at start-up and loaded from there. Nothing you type leaves this server.
01 / decision
Give them a real choice
Pick a demo or write your own. Questions use the same JSON shape as the Laya and Jev SDKs:
choice for labelled options, score for an ordered scale, noul for yes or no.
Choose a demo
02 / verdicts
Side by side
Each option shows one bar per model, in the order of the cards above; the bold bar is that model's top answer. A verdict is marked act when its confidence clears your threshold and defer when a person should decide.
Run a decision to see every model's answer here.
Raw JSON
03 / scoreboard
Every demo, ranked for agents
Runs every demo through every model, compares each answer with a reference answer (a careful human reading, not ground truth), highlights the better model on each metric and ranks them by the Agentic Use Score. A few dozen demos show behaviour, not accuracy; test on your own labelled data before trusting a threshold.
04 / reading the numbers
What these numbers do, and do not, mean
The Agentic Use Score
When an agent acts on an answer, the mistakes that hurt are the confident ones: the table it deletes, the payment it sends. A mistake the model defers costs only a quick human look. A correct answer it defers costs a little time.
So the score, from 0 to 100, rewards a model most for being right when it acts and for flagging its own mistakes.
High-stakes questions (irreversible actions, security, money, personal data) count twice, except in speed. Moving the threshold or changing the confidence measure changes the score, because together they decide when a model acts.
- 35% Right when it actsOf the answers confident enough to act on, how many were correct. A small allowance stops a model that acts only once or twice from scoring perfectly.
- 25% Flags its own mistakesOf the answers that were wrong, how many were unsure enough to go to a person instead.
- 15% Overall accuracyHow many answers match the reference, acted on or not.
- 15% Handles on its ownHow many decisions clear the threshold, so the agent doesn't need a person.
- 10% SpeedFull marks at 50 ms or faster, none at one second or slower. Agents make many decisions per task.
Two kinds of confidence
LightDec (both versions) and Jev report the probability of the top option.
Laya reports how concentrated the whole distribution is (1 − normalised entropy) for choice and score questions, and the larger of P(yes) and P(no) for yes/no.
The lab computes both measures for every model. Pick one with Confidence means.
Real local timing
Each model is timed on this server with a high-resolution clock.
The models run one after the other, so they never compete for the device.
Times depend on the hardware: a GPU is much faster than a CPU.
The models
LightDec_Arthur: byte-level Arthur model, about 12M parameters; the code that runs it ships with DecisionLab.
LightDec_V2: FalconDec on the Ettin-150M encoder, 2,048-token window, Apache-2.0.
Enterprise Reflux Laya V2.1: a fine-tune of Laya for ranking enterprise actions, by Mohamed Yasser; its model card marks it a research prototype, not for production.
Laya: by Convai Innovations, ModernBERT-large, Apache-2.0.
Models in the mounted models folder are added after these, marked "(local)": a LightDec folder has a falcondec_config.json, an Arthur folder a config.json and model.safetensors.
None of the models writes text. Each only ranks the options you give it.