1. Public benchmarks
Standard datasets run with fixed prompts, versions and decoding settings.


OUTAI will publish results only after the relevant model exists, the test procedure is documented and the evidence can be reproduced.
Standard datasets run with fixed prompts, versions and decoding settings.
Representative enterprise and regional tasks with controlled ground truth.
Prompt injection, ambiguity, unsupported claims and unsafe edge cases.
Expert review where automated scoring is insufficient.
| Model | Status | Scores | Evidence |
|---|---|---|---|
| OUTAI Base | Planned | Not published | Pending model selection |
| OUTAI Reason | Planned | Not published | Pending training |
| OUTAI Code | Planned | Not published | Pending training |
| OUTAI Medical | Research | Not published | Clinical validation required |
This website intentionally contains no fabricated benchmark claims. Verified results can replace these placeholders after formal evaluation.