Atomicwork and New Measure have introduced ITSMBench, a benchmark designed to evaluate how effectively AI models and agent frameworks perform real-world IT Service Management tasks. The open-source benchmark recreates enterprise service desk environments using 42 mocked software systems, nearly 1,800 database tables, and more than 2,000 REST endpoints, covering 89 L2 and L3 tasks across identity and access management, networking, security, infrastructure, devices, and engineering. Initial evaluations found clear differences among leading models in tool discovery, task execution, cost, and verification. Grok 4.5 achieved the highest tool-discovery rate at 83%, while Opus 5 recorded the strongest execution performance, completing 63.5% of tasks after identifying the appropriate tools. The benchmark also showed that agent frameworks can significantly influence accuracy and cost, with models sometimes stopping after finding plausible explanations instead of confirming root causes or completing follow-up actions.
Also Read: Deployable Energy Secures Solaris Energy Investment to Advance Deployable Microreactor Commercialization
“Frontier models are brilliant at writing code, but they are completely blind to the hidden security landmines inside enterprise workflows,” said Vijay Rayapati, co-founder and CEO of Atomicwork. “ITSMBench proves that blindly trusting an AI agent right now is an invitation for an enterprise security breach.” “Enterprise service management demands more than reasoning,” said Arushi Gandhi, CEO at New Measure. “One model fits all is seldom the right answer. Enterprises that want an AI-run service desk need to optimize for accuracy, cost, and time to resolution at the same time. Getting this right, thus, takes a constellation of orchestrated models. ITSMBench is our proposed framework for CIOs and IT leaders to evaluate which models/harnesses are best for them for their enterprise IT needs.”



