Impact
Exams don’t run plants.
So we built a benchmark that looks like the job: plant tasks, run in context, on real drawing sets. Everything on this page is measured, not claimed. The methodology is public, and the harness runs on any operator’s own drawings.
The exam, first
8× on the certification test.
On the PE certification test for engineering domain knowledge, our model scored 80.9%. GPT-4o scored 10.8%. DeepSeek RL 33b scored 20%. The 80.9% is our fifth training iteration.
| Model | PE certification test |
|---|---|
| Intuigence (fifth training iteration) | 80.9% |
| DeepSeek RL 33b | 20% |
| GPT-4o | 10.8% |
The benchmark that matters
Measured on the work itself.
The in-context plant-task benchmark runs the tasks a plant actually assigns: tracing lines across sheets, extracting what the drawings hold, answering engineering questions against a compiled estate. Real drawing sets, not curated samples.
Public methodology
The task definitions and scoring are published. Nothing on this page depends on taking our word for it.
Your drawings, not ours
The harness runs on any operator’s own drawing sets. That is the offer: run it on yours, and score us on your plant.
Judged by your engineers
The output is engineering work with citations attached. The people who can tell right from plausible are your own.
“It did in an afternoon what my team had backlogged for a quarter.”
Senior process engineer, Fortune-100 refiner
The numbers
“I see IntuigenceAI as the next generation of industrial technology.”
Doug Raven, Former Senior Engineer, Saudi Aramco