创建

登录 ReadmeX

登录后可以加入社区、发帖、投票和聊天。

还没有账号?

资讯

AWS DevOps Agent 自动修复工作流教程

AI 总结

AWS 机器学习博客发布教程,演示如何在 AWS DevOps Agent 完成事故调查后自动执行修复。方案由 Amazon EventBridge 接收调查完成事件并触发 AWS Lambda Durable Functions,再由 Amazon Bedrock 在预置的 Lambda 工具白名单内提出修复动作;只读操作自动执行,会改动基础设施状态的操作则暂停等待人工批准。示例场景为修复一个 Lambda 函数超时时间配置过短的问题。

为什么重要:它展示了一种“调查—预校验—单次审批—修复”的人机协同模式,可在保留人工控制的前提下缩短平均修复时间(MTTR),但目前仍是 AWS 官方博客中的示例实现,需要用户自行部署。

AWSAWS DevOps AgentAWS Lambda Durable Functions

9
来源原文AWS Machine Learning Blog · 约 11 分钟读完

Reducing the time between incident detection, investigation, and remediation is a critical priority for organizations running production workloads on AWS. When an issue arises, on-call engineers often need to quickly diagnose the problem across application components, identify the root cause, and apply the fix, often in the middle of the night.

AWS DevOps Agent, an AI powered agent that autonomously triages incidents all day based on correlated metrics, logs, and application topology, addresses the first part of this priority by providing root cause analysis (RCA) and recommended actions for resolution. However, to retain control and help prevent unintended changes, organizations typically keep their observability agents, including AWS DevOps Agent, in an observe-and-report mode, where the agent diagnoses issues but doesn’t modify production resources directly. In this post, we demonstrate how to use AWS Lambda Durable Functions, a capability of AWS Lambda, Amazon EventBridge, and Amazon Bedrock to create an automated remediation workflow that complements AWS DevOps Agent to complete the issue resolution step. This workflow transforms investigation summaries into pre-validated fixes ready for single approval action, helping you reduce mean time to resolution (MTTR) and free your on-call engineers from repetitive diagnostic work.

Solution overview

With AWS Lambda Durable Functions, you can build resilient multi-step applications and AI workflows that can run for up to one year without requiring you to manage additional infrastructure or write custom state management and error handling code. These functions automatically checkpoint progress, suspend execution during long-running tasks, and recover from failures while maintaining reliable progress despite interruptions.

The following diagram illustrates the solution architecture.

Figure 1: Automated remediation workflow from an AWS DevOps Agent investigation through Amazon EventBridge, AWS Lambda, and Amazon Bedrock, with optional human approval before changes reach the infrastructure

The workflow consists of the following steps:

  1. AWS DevOps Agent completes an incident investigation and emits an event containing symptoms, findings, and root cause analysis.
  2. Amazon EventBridge receives the investigation completion event and triggers the devops-agent-trigger function with the investigation content.
  3. The Lambda function packages the investigation summary and invokes the devops-agent-remediation-durable durable function.
  4. The durable function sends the investigation context to Amazon Bedrock, which analyzes the findings and looks for applicable remediations.
  5. Amazon Bedrock identifies and lists the available remediation tools from a curated allowlist of approved Lambda functions: devops-agent-lambda-tool.
  6. Amazon Bedrock proposes specific remediation actions based on the investigation findings and the available tools.
  7. For read-only actions, the durable function runs the remediation tools autonomously. For infrastructure changes, the workflow suspends and waits for human approval before proceeding.
  8. After approval, the durable function applies the remediation actions to the infrastructure using the selected tools.

The durable function runs as an agentic loop, iteratively calling Amazon Bedrock, executing approved tools, and feeding results back into the conversation until the remediation is complete. To keep automated actions safe and auditable, the orchestrator enforces a curated allowlist of remediation tools. Each tool is a purpose-built Lambda function that performs a specific, well-scoped action, such as reading a Lambda function configuration or updating an AWS Identity and Access Management (IAM) policy statement. Amazon Bedrock can only select and invoke tools from this approved set, which keeps the scope of automated actions controlled. The workflow further distinguishes between read-only operations and mutating operations. Read-only tools run autonomously without human intervention. Mutating actions that would modify infrastructure state cause the durable function to suspend execution and wait for human approval. This is where AWS Lambda Durable Functions provide a key advantage. The function checkpoints its progress and pauses for minutes, hours, or even days without consuming compute resources, then resumes exactly where it left off after it receives the approval signal. By the time the on-call engineer engages, the system has already gathered relevant configurations, correlated the root cause with available remediation actions, and prepared a set of pre-validated changes ready for one-click approval. The current implementation uses an approve or reject signal. Because the callback accepts an arbitrary JSON payload, you can extend the approval to carry parameter overrides or reviewer observations. These can be fed back into the Bedrock conversation to refine the proposed remediation before execution.

In the following sections, we walk through the implementation details, including the Amazon EventBridge rule configuration and the durable function orchestration logic. We then deploy the solution using the AWS Cloud Development Kit (AWS CDK).

Prerequisites

Before deploying this solution, verify that you have the following prerequisites:

  • The AWS Command Line Interface (AWS CLI) installed and configured.
  • Python 3.14 or later.
  • The AWS CDK installed.
  • An active AWS DevOps Agent space.
  • (Optional) Kiro with the Agent Toolkit for AWS. The Agent Toolkit gives Kiro secure access to AWS APIs through a managed MCP Server with IAM-based access controls. If you use Kiro, the incident simulation, deployment, and cleanup steps in this post can be completed with natural language prompts instead of running CLI commands manually. To set it up, add the AWS MCP Server to ~/.kiro/settings/mcp.json (setup instructions). The repository includes a Kiro rules and agents file that gives Kiro the project context, deployment sequence, and safety conventions automatically.

Simulate the incident

To demonstrate the end-to-end workflow, we simulate a common scenario: a Lambda function that exceeds its configured timeout. This gives AWS DevOps Agent a real incident to investigate and sets the remediation workflow in motion.

To keep the focus on the remediation solution itself, the steps to create and invoke this test function are kept in the repository. It includes a ready-to-use devops-agent-timeout function and step-by-step instructions to deploy it, invoke it, and confirm the timeout error in Amazon CloudWatch Logs. For the full walkthrough, see the “Simulate the incident” section of the README.

After the function is deployed and has produced at least one timeout error, you’re ready to start an investigation with AWS DevOps Agent.

Deploy the solution using the AWS CDK

Complete the following steps to deploy the remaining solution resources:

Kiro: If you have Kiro with the Agent Toolkit for AWS configured (see Prerequisites), open the cloned repository in Kiro and ask: “Set up the Python environment and deploy the CDK stack. Show me what resources will be created before deploying.” Kiro reads the project rules from the repository, sets up the virtual environment, installs dependencies, and shows you the planned resources before deploying. It confirms each infrastructure change before executing, following the same human-in-the-loop pattern that the remediation solution itself uses. To deploy manually, follow these steps.

  1. Clone the AWS CDK code hosted on GitHub:
    $ git clone https://github.com/aws-samples/sample-automate-remediation-post-devops-agent-investigation.git
  2. Navigate to the directory sample-automate-remediation-post-devops-agent-investigation:
    $ cd sample-automate-remediation-post-devops-agent-investigation
  3. Bootstrap the AWS CDK. This is required the first time you use the AWS CDK in a specific AWS environment (a combination of an AWS account and AWS Region).
    $ cdk bootstrap
  4. Deploy the stack:
    $ cdk deploy

The AWS CDK automatically provisions and configures the following resources:

  • Three Lambda functions:
    • devops-agent-trigger.
    • devops-agent-remediation-durable.
    • devops-agent-lambda-tool.
  • Amazon EventBridge rule.

The AWS CDK automatically handles the IAM permissions using least-privilege principles and AWS security best practices. For example, Amazon EventBridge is granted lambda:InvokeFunction permissions for the devops-agent-trigger function. The stack grants the aidevops:ListJournalRecords permission to the devops-agent-trigger function so it can fetch investigation summaries from the AWS DevOps Agent journal. It also grants the bedrock:InvokeModel permission to the devops-agent-remediation-durable function so it can invoke Amazon Bedrock.

Validate the solution

With the remediation stack deployed and the devops-agent-timeout function failing with timeout errors, we can now walk through the end-to-end workflow.

Start an investigation with AWS DevOps Agent

Open the AWS DevOps Agent console, navigate to your agent space, and ask: “What is happening with the devops-agent-timeout function?”

Figure 2: Starting an investigation from the AWS DevOps Agent console

The investigation starts and takes a few minutes to complete. During this time, AWS DevOps Agent autonomously correlates CloudWatch metrics, logs, and the function’s configuration to determine the root cause.

Figure 3: AWS DevOps Agent correlating signals during the investigation

After the investigation completes, AWS DevOps Agent presents the root cause analysis, identifying that the function timeout is insufficient for the workload.

Figure 4: The root cause analysis identifying the insufficient function timeout

Verify the trigger Lambda execution

The investigation completion emits an Investigation Completed event to Amazon EventBridge.

The rule triggers the devops-agent-trigger Lambda function, which fetches the investigation summary from the AWS DevOps Agent journal. In the /aws/lambda/devops-agent-trigger CloudWatch log group, you can see the parsed summary that is sent to the devops-agent-remediation-durable durable function, including symptoms, root causes, contributing causes, and investigation gaps.

Figure 5: The parsed investigation summary in the trigger function CloudWatch log group

Monitor the durable function execution

Navigate to the Lambda console, open the devops-agent-remediation-durable function, and choose the Durable executions tab. Choose the new execution to inspect its checkpointed steps.

Figure 6: The durable function execution and its checkpointed steps in the Lambda console

The durable orchestrator begins its agentic loop by sending the investigation context to Amazon Bedrock. In the first Bedrock call, the model analyzes the investigation summary and determines that it needs to inspect the current function configuration before proposing a fix. It selects the lambda_get_function_configuration tool from the allowlist. Because this is a read-only operation, it runs autonomously without requiring human approval. The step result shows the current configuration of the devops-agent-timeout function, confirming a timeout value of 3 seconds.

Figure 7: The read-only tool call returning the current 3-second timeout configuration

Amazon Bedrock proposes the remediation

With the current configuration confirmed, Amazon Bedrock proceeds to the next iteration. It reasons that the 3-second timeout is the root cause of the failures and proposes increasing it to 30 seconds. The Amazon Bedrock response contains both the reasoning and the tool call:

{
  ....
  },
  "output": {
    "message": {
      "role": "assistant",
      "content": [
        {
          "text": "Now I can see the current configuration confirms the issue - the function has a 3-second timeout. Given that this appears to be a test function for timeout scenarios (based on the name \"devops-agent-timeout\" and its association with \"DevOps Agent Test Infrastructure\"), I'll increase the timeout to a reasonable value that should allow the function to complete successfully. I'll set it to 30 seconds, which is a common timeout for Lambda functions that need more execution time."
        },
        {
          "toolUse": {
            "toolUseId": "tooluse_lcGHbWDxng5HiLpmhWagDS",
            "name": "lambda_update_function_configuration",
            "input": {
              "FunctionName": "devops-agent-timeout",
              "Timeout": 30
            },
            "type": "tool_use"
          }
        }
        ...
      }

Because lambda_update_function_configuration is a mutating action, the durable function suspends execution and waits for human approval.

cloudwatch logs output of lambda durable function for approval request Figure 8: cloudwatch logs output of lambda durable function for approval request

Important: The investigation_summary sent to Amazon Bedrock, and the remediation it proposes, are AI-generated and should always be reviewed before approval. The human approval gate is the security control: the approver must inspect the full tool parameters (for example, the exact FunctionName and Timeout in a lambda_update_function_configuration call) and confirm the change is correct.

Using the AWS CLI:

$ aws lambda send-durable-execution-callback-success \
    --callback-id <callback-id> \
    --cli-binary-format raw-in-base64-out \
    --result '{"approved": true}'

Using the AWS Console:

Navigate to the durable execution, select the pending callback, and choose Send success to confirm:

Figure 9: Approving the remediation by choosing Send success in the Lambda console

In the input field, enter {'approved': true} and confirm.

Figure 10: Entering the approval payload to confirm the callback

Verify the fix

After approval, the durable function resumes, invokes the tool Lambda to update the configuration, and Amazon Bedrock confirms the remediation is complete. The updated devops-agent-timeout function now shows the new timeout value:

Figure 11: The devops-agent-timeout function configuration updated to a 30-second timeout

The final step output (bedrock-call-4) confirms the successful remediation:

{
  "EventType": "StepSucceeded",
  "Name": "bedrock-call-4",
  ....
  },
  "output": {
    "message": {
      "role": "assistant",
      "content": [
        {
          "text": "## Remediation Complete
          **Issue Identified:** The Lambda function `devops-agent-timeout` in eu-west-1 was configured with a 3-second timeout that was insufficient for its execution time, causing timeout errors at exactly 3000ms.
          **Action Taken:** Successfully updated the Lambda function configuration to increase the timeout from 3 seconds to 30 seconds.
          **Verification:** Confirmed the configuration change was applied successfully. The function now has:
          - **Timeout:** 30 seconds (increased from 3 seconds)
          - **Status:** Successful update completion
          - **LastModified:** 2026-05-22T11:11:28.000+0000
          This remediation should resolve the timeout errors by providing the function with adequate time to complete its execution. The 30-second timeout provides a 10x increase from the original 3-second limit, which should be sufficient for most operations while still maintaining reasonable execution bounds for a Lambda function."
        }
      ]
    }
  },
  "stopReason": "end_turn",
  ...
}

This entire cycle, from incident detection to automated fix, required only a single approval action from the engineer. The system handled diagnosis, configuration retrieval, remediation proposal, and execution autonomously.

Clean up

Clean up the resources you created by completing the following steps:

Kiro: If you use Kiro with the Agent Toolkit for AWS, ask: “Clean up all resources from the DevOps Agent remediation demo: destroy the CDK stack, delete the devops-agent-timeout test function, its IAM role, and its CloudWatch log group.” Kiro removes resources in the correct order, confirming each destructive action before proceeding. To clean up manually, follow these steps.

  1. Delete the AWS CDK resources:
    $ cdk destroy
  2. Manually delete the devops-agent-timeout function that simulates the incident:
    $ aws lambda delete-function --function-name devops-agent-timeout
  3. Manually delete the IAM role and the CloudWatch log group of the devops-agent-timeout function:
    $ aws iam detach-role-policy \
        --role-name devops-agent-timeout-role \
        --policy-arn arn:aws:iam::aws:policy/service-role/AWSLambdaBasicExecutionRole
    $ aws iam delete-role --role-name devops-agent-timeout-role
    $ aws logs delete-log-group --log-group-name /aws/lambda/devops-agent-timeout

Conclusion

This post demonstrated how you can automate issue remediation by using Lambda Durable Functions, Amazon EventBridge, and Bedrock in conjunction with the DevOps Agent. The solution picks up where AWS DevOps Agent leaves off, transforming investigation summaries into actionable remediation steps that run with human-in-the-loop approval. This approach reduces mean time to resolution because, by the time the on-call engineer engages, the system has already diagnosed the issue, gathered current configurations, and prepared a ready-to-approve fix. Safety remains central to the design: the allowlist restricts Amazon Bedrock to only invoke pre-approved tools, and the human approval gate helps prevent unintended changes from reaching production without explicit authorization. The architecture is also inherently extensible. Adding new remediation capabilities requires only configuration updates to the tool registry, not code changes to the orchestrator. And because AWS Lambda Durable Functions suspend without consuming compute resources during the approval wait, the solution remains cost-efficient even when approval cycles span hours or days.

To get started using this solution, download the complete AWS CDK template from the GitHub repository, and follow the steps in this post to deploy the solution in your environment.

We would love to hear from you. Share your experience implementing this solution, ask questions, or suggest improvements in the comments. You can also join the AWS Community Builders program to connect with other builders and share your serverless architecture patterns.


About the author

阅读原始来源 →

前因后果

  1. Qlik Answers:在 Amazon Bedrock 上规模化落地企业级 AIAWS Machine Learning Blog · Amazon Bedrock
  2. AWS 六周培养计划:让非工程人员亲手构建 AI 智能体AWS Machine Learning Blog · Amazon Bedrock
  3. Cornerstone用Bedrock多智能体系统将数据库诊断提速78%AWS Machine Learning Blog · Amazon Bedrock
  4. AWS 解读 ISO/IEC 42005 影响评估标准AWS Machine Learning Blog · AWS
  5. AWS 教程:用 AgentCore 与 Nova Sonic 构建语音旅行助手AWS Machine Learning Blog · AWS
  6. Anthropic扩大CVP:三级资安权限开放ClaudeAnthropic · News · Amazon Bedrock

评论

我用过:分享经验 我怎么看:发表观点
你觉得这条新闻有多重要?还没有评分

还没有评论,来说说你的看法。