Laravel AI SDK 1.0 shipped with something I had not seen in a Laravel package before: classification as a capability of its own, built for a new kind of model. Not text generation with a JSON schema bolted on, but a model that only answers typed questions. TypeSafe call it a System One model, and theirs is called Jev. I wanted to know what it is actually like to build with, so I tried it on a job where a confident wrong answer would really matter: marking Year 6 science exams.
The result is a small Laravel app and an interactive showcase you can play with in the browser. This post is what I found along the way.
The first five minutes
The quickest way in is a string macro. Str::decide() asks Jev one yes-or-no question and gives you a boolean back, with a threshold for how sure it has to be:
use Illuminate\Support\Str;
$attempted = Str::decide(
'the sun heats the water and it evaperates into the air',
'Is this a genuine attempt at a science question?',
threshold: 0.8,
provider: 'openrouter',
model: 'typesafe/jev-1.13',
);
That is the whole call. No prompt engineering, no "respond only with true or false", no parsing. The pupil spelt "evaporates" wrong and it does not matter: the question is whether this is a genuine attempt, and the answer comes back as a probability that the macro compares against 0.8.
For anything richer you use Classification::of() with typed questions. Marking needs a number, so the first real question was a Score:
use Laravel\Ai\Classification;
use Laravel\Ai\Classification\Score;
$response = Classification::of([
'question' => 'Why does a puddle disappear on a sunny day?',
'mark_scheme' => 'One mark for evaporates. One mark for heat turns liquid water into water vapour.',
'pupil_answer' => 'the sun heats the water and it evaperates into the air',
])
->question('mark', new Score('How many marks does the answer earn? Ignore spelling.', [
'0 marks: meets no point in the mark scheme',
'1 mark: meets 1 of the 2 points',
'2 marks: fully meets the mark scheme',
]))
->classify(provider: 'openrouter', model: 'typesafe/jev-1.13');
$mark = $response->answer('mark');
$mark->score; // a position on the scale, such as 1.99
$mark->confidence; // how sure Jev is, such as 0.97
What comes back
This was the moment it clicked. A Score answer is not "2". It is a position on the scale, such as 1.99, plus a confidence and the full probability spread across every level. You decide how to round it, and you decide what a confidence of 0.70 means for your app. Jev never pretends to be certain; it tells you how certain it is, every time.
There are three question types in total, and they cover a surprising amount of ground:
- Noul: yes or no. Jev returns the probability of yes, so 0.04 is a confident no
- Choice: pick one option. Jev returns the option, a confidence and the spread across all of them
- Score: a position on an ordered scale, lowest first
Asking three questions about every answer
One question was not enough to mark fairly. A blank answer and a wrong answer both score 0, but they need different handling, and a misconception ("light comes out of our eyes") is worth flagging to a teacher in a way that a simple slip is not. So every pupil answer gets three questions in one request: mark (a Score), makes_sense (a Noul) and error_type (a Choice).
What surprised me is that the criteria are the prompt. Each option is described in plain words, and that wording is where all the judgement lives:
{
"type": "choice",
"instructions": "Which best describes the pupil answer?",
"criteria": {
"correct": "Fully meets the mark scheme; spelling mistakes are ignored",
"partially_correct": "Meets part of the mark scheme but misses at least one required point",
"misconception": "Attempts the question but shows a scientific misunderstanding",
"incomplete": "Starts a valid answer but stops before making the point",
"off_topic": "Does not address the question that was asked"
}
}
I ended up keeping those descriptions on a PHP enum, so the labels a teacher sees in the app and the criteria Jev is given come from exactly the same place. Change the wording and you change how answers are classified, which is easy to reason about in a way prompt tweaks rarely are.
Turning typed answers into marks
Because every answer arrives with a confidence, deciding what to do with it is ordinary code. That was the part I enjoyed most: no model in the loop, just a plain PHP class with its thresholds in config.
public function routeAnswer(MarkingDecision $decision, int $maxMarks): AnswerRoute
{
$reasons = [];
if (! $decision->makesSense()->isTrue()) {
$reasons[] = 'Jev thinks the answer does not make sense';
}
$confidences = [
MarkingQuestions::MARK => $decision->mark()->confidence(),
MarkingQuestions::MAKES_SENSE => $decision->makesSense()->confidence(),
MarkingQuestions::ERROR_TYPE => $decision->errorType()->confidence(),
];
foreach ($confidences as $question => $confidence) {
$threshold = $this->thresholds[$question];
if ($confidence < $threshold) {
$reasons[] = sprintf('%s confidence %.2f is below %.2f', $question, $confidence, $threshold);
}
}
return new AnswerRoute(
mark: $decision->markOutOf($maxMarks),
makesSense: $decision->makesSense()->isTrue(),
errorType: $decision->errorTypeEnum(),
reasons: $reasons,
);
}
Three rules came out of trying it on real-looking answers:
- if every answer makes sense and every confidence clears its threshold, the paper is marked and the result goes back to the teacher
- if any answer does not make sense, or any confidence is below its threshold, that answer goes to a human marker with Jev's answers shown as suggestions, and the result waits
- if the total lands within two marks of the pass mark, a human confirms it however confident Jev was, because that is exactly where a one-mark mistake changes a pupil's result
The consequence is the bit worth taking away: the same answers and the same thresholds always give the same outcome. You can explain to a teacher exactly why a paper came back or did not, down to "mark confidence 0.70 is below 0.75". The showcase lets you drag those thresholds yourself and watch the decision flip.
Where the SDK stopped fitting
The SDK got me started fast, and I kept it, but not for Jev in the end. Reading its source turned up a few things worth knowing before you build on it:
- the OpenRouter driver calls
/api/alpha/decisions; the TypeSafe driver calls/v1/systemone. Both share one request and response shape - the response object keeps token counts and the model name but drops the cost and the request id that OpenRouter returns
- the default model is the latest alias. For marking I pinned
typesafe/jev-1.13, because a silent model upgrade would move marks between one week and the next
I wanted a log of every call with its cost, so the marking path uses a thin client of my own behind a DecisionClient interface. It is one POST over Laravel's Http client. The SDK stayed in the app for the parts where a generative model genuinely helps.
Where a generative model still earns its place
Jev does not write, and for marking that is a feature. But teachers need words too, so I added three things with the SDK's agent() helper, switchable between Claude and OpenAI by config:
- a sentence or two of feedback for the pupil on each answer that lost marks
- a short report comment, whose result line ("Ava: 19/20, pass") is written by code from the stored marks
- a word-for-word transcript of a scanned handwritten paper, which the teacher checks before Jev sees it, because a transcriber that tidies "evaperates" would change the pupil's mark
The rule I held to throughout: Jev decides, code acts, a human has the final say, and the generative model only writes about decisions that are already final.
If you want to try it
- start with
Str::decide()on something you already route by hand, and look at the probabilities before you pick a threshold - write the criteria as carefully as you would write a mark scheme; they are the prompt
- keep thresholds in config and the routing in plain code, and give "not sure" a human destination
- pin the model version, and if you need cost per call, check what the response object actually keeps
- build an offline fake so you can work without a key; the repo has one built from hand-labelled answers
I have now run it against live Jev, through TypeSafe's own API. A full-marks paper came back 20/20, with every answer marked correct at 0.99 confidence or higher, in around 370 ms a call. The first surprise was a misconception: "Our eyes shine light onto the book" scored 0 with 0.99 confidence and was classified as a misconception at 1.00, so it was marked automatically, where my offline fake had sent it to a human. The mark is right, but the teacher no longer sees that misconception in the queue.
So I ran the full evaluation: all 200 hand-labelled answers through live Jev, one call each, in about 70 seconds and for roughly half a cent. Jev matched my mark exactly on 197 of them (98.5%) and was never more than one mark out. It caught every blank, joke and nonsense answer, and agreed on the error type 95% of the time, including every misconception. With the current thresholds, 83% of answers were marked automatically, and no paper went back to a teacher with the wrong pass or fail without a human checking it first.
All three disagreements were on the same question, how a fossil forms, where Jev gave 1 mark out of 2 and my labels gave 2. Two of them came back below the confidence threshold and went to a human, which is the design doing its job. Reading the answers again, I think Jev had a point on those two: "over a very long time it turns to rock" does not quite say the sediment becomes rock, which is what the mark scheme asks for. The third answer did say it, and Jev was wrong. That is the kind of finding an evaluation run is for: it tells you where your labels are generous, where the mark scheme is ambiguous, and where the model needs a human behind it.
Everything else, including the seeded class, sample papers and a handwritten scan to transcribe, works without a key.
The code is on GitHub and the showcase is on GitHub Pages. If you are weighing up a decision model for routing or triage work in a Laravel app and want a second pair of eyes, get in touch via the contact section.