Amazon Bedrock AgentCore
A Lambda subscribed to the AgentCore runtime log group. With message content capture on, ADOT writes every GenAI message as its own log record; a subscription filter forwards them, the Lambda rebuilds one evaluation unit per model call and sends the user prompt, tool arguments, tool results and model output to Alice. Verdicts land in CloudWatch and, optionally, as spans inside the agent trace. Nothing in the agent changes and nothing is blocked.
Status: Generally available · Evaluates: Observability · Vendor: AWS
Input, tool arguments, tool results and output, reconstructed from telemetry after the fact.
Setup
sam build && sam deploy
Configuration: AliceApiKeySecretArn, AgentConfig (service.name=app_id), SourceLogGroupName, EmitTraceSpans
- Set
OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=trueon the agent runtime. - Store the Alice API key as an AgentCore Identity credential provider (or a Secrets Manager secret).
- Deploy the SAM stack against the runtime log group with a strict
service.name=app_idallowlist. - Read verdicts in the Lambda log group, or enable X-Ray spans and filter on
alice_action.
Example
deploy the monitor (SAM)
RUNTIME_ID=$(agentcore status --json | jq -r '.deployedState.targets.default.resources.runtimes.MyAgent.runtimeId')
SECRET_ARN=$(aws bedrock-agentcore-control get-api-key-credential-provider --name AliceApiKey --query 'apiKeySecretArn.secretArn' --output text)
cd monitor
sam build
sam deploy --stack-name alice-agentcore-telemetry-monitor --resolve-s3 --capabilities CAPABILITY_IAM --no-confirm-changeset --parameter-overrides "AliceApiKeySecretArn=$SECRET_ARN" "AgentConfig=<service.name>=<ALICE_APP_ID>" "SourceLogGroupName=/aws/bedrock-agentcore/runtimes/$RUNTIME_ID-DEFAULT" EmitTraceSpans=true
handler.py: client, context and per-span evaluation
def _build_client() -> WonderFenceV2Client:
"""Construct the WonderFence client.
The SDK resolves its base URL as ``base_url or os.environ["ALICE_URL_OVERRIDE"]``
and falls back to its built-in default only when that env var is *absent*. Our
template always sets ALICE_URL_OVERRIDE (empty by default), and an empty string
would shadow the SDK default and fail validation — so resolve an empty override
to ``None`` and pop it, letting the SDK apply its default.
"""
url_override = os.environ.pop("ALICE_URL_OVERRIDE", "").strip() or None
return WonderFenceV2Client(api_key=_load_api_key(), base_url=url_override)
# Initialize at module level to persist across warm invocations.
# V2 client takes app_id per request, so a single client serves all agents.
client = _build_client()
AGENT_CONFIG = _load_agent_config()
# Opt-in: also emit each evaluation as an X-Ray subsegment so it shows up in the
# AgentCore / X-Ray (Transaction Search) trace tree (see trace_emitter).
ADD_TO_AGENT_TRACE_SPANS = (
os.environ.get("ADD_TO_AGENT_TRACE_SPANS", "false").strip().lower() == "true"
)
def _truncate(text: str) -> str:
"""Truncate text to MAX_TEXT_LENGTH to avoid API timeouts."""
if len(text) > MAX_TEXT_LENGTH:
return text[:MAX_TEXT_LENGTH]
return text
def _build_context(span: ParsedSpan) -> AnalysisContext:
return AnalysisContext(
session_id=span.trace_id,
provider=span.gen_ai_system,
platform="aws-agentcore",
)
def _evaluate_span(span: ParsedSpan, app_id: str) -> list[dict]:
"""Run evaluations on a parsed span against the given Alice app_id."""
context = _build_context(span)
evaluations = []
eval_tasks = []
if span.input_text:
eval_tasks.append(("prompt", "input_text", span.input_text))
if span.tool_arguments:
eval_tasks.append(("prompt", "tool_arguments", span.tool_arguments))
if span.output_text:
eval_tasks.append(("response", "output_text", span.output_text))
if span.tool_result:
eval_tasks.append(("response", "tool_result", span.tool_result))
for eval_type, field_name, text in eval_tasks:
try:
truncated = _truncate(text)
if eval_type == "prompt":
result = client.evaluate_prompt_sync(app_id, context, prompt=truncated)
else:
result = client.evaluate_response_sync(app_id, context, response=truncated)
evaluations.append(
{
"type": field_name,
"action": result.action.value
if hasattr(result.action, "value")
else str(result.action),
"detections": [{"type": d.type, "score": d.score} for d in result.detections],
"correlation_id": result.correlation_id,
}
)
except Exception:
logger.exception(
"Evaluation failed for %s on span %s/%s",
field_name,
span.trace_id,
span.span_id,
)
evaluations.append(
{
"type": field_name,
"action": "ERROR",
"error": "evaluation_failed",
}
)
return evaluations
query the verdicts
aws logs filter-log-events --log-group-name "/aws/lambda/alice-agentcore-telemetry-monitor" --filter-pattern '"alice-agentcore-monitor"'
Good to know
Observe only: an Alice outage means no verdicts, the agent is unaffected. X-Ray spans need CloudWatch Transaction Search. One stack per runtime; one Lambda routes many agents via AgentConfig. Text is truncated to 10 KB per field.