Block AI Agent Regressions Before They Ship | SAO Pre-Push Eval Gate Demo
Every engineering team has unit tests. They tell you the code still works. They tell you nothing about what the model started saying. This demo wires a single eval gate script into a git pre-push hook, so Splunk Agent Observability scores every agent's output before the push is allowed through. Luna, an on-premise small language model, runs as a synchronous judge against fixed thresholds. Fail one, and the push is blocked.