MikhbarMIKHBAR
Artificial Intelligence

Cornerstone Cuts Database Diagnosis by 78% Using Amazon Bedrock

Cornerstone OnDemand transformed its data operations from reactive firefighting into proactive automation by developing a multi-agent system called Orion AI.

Cornerstone Cuts Database Diagnosis by 78% Using Amazon Bedrock

Transforming Data Operations with Orion AI

Cornerstone OnDemand, Inc. (Cornerstone) serves 140 million users across 186 countries as a global leader in workforce readiness solutions. To tackle complex database operations, the company engineered a multi-agent AI system named Orion AI. The platform is designed to shift enterprise workflows from reactive firefighting into proactive, self-orchestrating processes by coordinating specialized agents via AWS infrastructure.

Before the introduction of Orion AI, the Enterprise DataOps team faced heavy friction during investigations. Database performance probes required approximately 45 minutes per incident across multiple tools and system views, while manual lifecycle workflows demanded ten or more distinct steps. Additionally, reporting between site reliability engineering and data teams suffered from a 15-minute lag, and redundant alerts created heavy noise.

Architecture and Design of the Multi-Agent System

Orion AI utilizes a hub-and-spoke topology where a meta-orchestrator delegates work to specialized child agents. Engineers interact with the system through a web application deployed as containerized services on Amazon Elastic Container Service (Amazon ECS). The architecture relies on managed access to foundation models and custom orchestration frameworks.

The core operational framework is built using Amazon Bedrock alongside Strands Agents, an open-source agent orchestration framework from AWS. Two primary tenets guided the development process: strict data privacy enforcement using AWS controls under the shared responsibility model, and deep integration with existing operational tooling.

Measurable Efficiency Gains and Speed Reductions

The implementation yielded significant improvements in diagnosis speed, accuracy, and team coordination. Previously, engineers spent 45 minutes manually connecting to SQL Server instances, querying system views for blocking chains, and cross-referencing logs. Orion AI reduced this database diagnosis timeline to just 10 minutes, representing a 78% reduction in troubleshooting duration.

Beyond diagnostics, the system carries issues through full remediation by identifying root causes, recommending fixes, and automatically populating Jira tickets assigned to the correct on-call engineer. Tasks requiring over ten manual steps now run as a single natural language interaction where the AI queries systems, validates results, and returns unified responses.

Architecture diagram of Orion AI within AWS Cloud. Web Application Users and external integrations (chat, issue tracking, incident management, and metrics dashboards) connect to a TaskExecutor agent running on Amazon ECS
Image related to the report from AWS Machine Learning Blog · Source: AWS Machine Learning Blog

Mitigating Alert Fatigue and Enhancing Visibility

By establishing continuous cross-system visibility, Orion AI effectively removed the historical 15-minute reporting lag between site reliability engineers and data teams. Periodic manual check-ins were replaced with automated real-time tracking.

Furthermore, the system achieved a median 65% reduction in redundant alerts through intelligent deduplication, threshold filtering, and cross-signal correlation. Out of every ten alerts previously generated, only three to four now reach engineers, routing directly to responsible teams without intermediate friction.

Reusable Design Principles for Engineering Teams

Cornerstone's engineering group outlined three core design decisions that other teams can reuse when building multi-agent projects. First, agents should be split by domain rather than task complexity, giving each agent a narrow set of tools to focus context windows and improve tool selection accuracy.

Second, teams should default to keyword routing for predictable requests while falling back to semantic search for ambiguous prompts, balancing low latency with high correctness. Third, conversational memory should be strictly scoped to sessions and bypassed for live metrics to prevent stale operational data from contaminating real-time diagnostics.

Sources

Continue chronologically

Related entity coverage