Base-model bake-off
Compare candidate open-weight models under the same inference and evaluation conditions.


OUTAI research begins with measurable questions: which licensed base, which training intervention, which data and which evaluation actually improve the target use case?
Compare candidate open-weight models under the same inference and evaluation conditions.
Measure Modern Standard Arabic, Gulf dialects, code-switching and enterprise terminology separately.
Test claim extraction, evidence checking, judge reliability and escalation thresholds.
Quantisation, routing, caching and private inference across server and edge targets.
A research claim must identify the model version, dataset, licensing status, method, hardware, evaluation procedure, limitations and reproducibility artefacts.