Skip to main content

Kong AI Gateway

Alice runs on a Kong AI Gateway through Kong's built-in ai-custom-guardrail plugin, a declarative HTTP callout: Kong sends the text it selected to Alice and enforces the verdict that comes back. Nothing custom runs in the gateway, so the whole integration is one decK config block. With request_body sent, Alice reads the conversation itself and screens the newest turn together with its tool calls and tool results.

Status: Beta · Evaluates: Prompts, Responses, Tool calls · Vendor: Kong

Prompts (INPUT) and responses (OUTPUT); with request_body, the newest turn and its tool calls and tool results.

Setup​

deck gateway sync kong.example.yaml

Configuration: DECK_ALICE_API_KEY, consumer custom_id = Application ID, guarding_mode, text_source, stop_on_error

  1. Requires Kong Gateway Enterprise 3.14 or later and an ai-proxy route for the guardrail to run around.
  2. Put key-auth in front of the route (key name x-api-key) and configure no anonymous consumer.
  3. Create one Kong consumer per application and set its Custom ID to the application's Application ID in Alice. That is how each request names whose policies apply.
  4. Add the ai-custom-guardrail block below, pointing at https://api.alice.io/v2/evaluate/kong with your Alice API key, and apply it with deck gateway sync.

Example​

kong.example.yaml (decK)

# Alice guardrails on a Kong AI Gateway, as a decK declarative config.
#
# Nothing custom runs in the gateway: `ai-custom-guardrail` is Kong's own plugin (Enterprise, 3.14+)
# and this file is the whole integration. Apply with:
#
# deck gateway sync kong.example.yaml
#
# Requires DECK_ALICE_API_KEY in the environment.
_format_version: '3.0'

services:
- name: openai
url: https://api.openai.com
routes:
- name: chat
paths:
- /chat
plugins:
# Authenticates every caller as a consumer — the consumer is what names the application.
# `x-api-key` is the header Anthropic SDKs and Claude Code send. Configure no `anonymous`
# consumer: with one, an unauthenticated call reaches the guardrail as that consumer.
- name: key-auth
config:
key_names:
- x-api-key
hide_credentials: true

# The AI Gateway proxy itself. The guardrail below runs around it, so it must be present.
- name: ai-proxy
config:
route_type: llm/v1/chat
auth:
header_name: Authorization
header_value: Bearer ${{ env "DECK_OPENAI_API_KEY" }}
model:
provider: openai
name: gpt-4o

- name: ai-custom-guardrail
config:
# Screen both halves of the turn. INPUT alone leaves the model's replies unscreened.
guarding_mode: BOTH

# What Kong sends as `$(content)`. `last_message` screens the newest turn only —
# the cheapest setting, and enough when every turn passes through this gateway.
# `concatenate_all_content` sends the whole conversation, which catches a violation
# split across turns at the cost of re-screening text earlier calls already saw.
text_source: last_message

# Alice answers well inside this. Kong abandons the callout when it expires and
# `stop_on_error` decides what that means for the caller.
timeout: 10000

# Fail closed: an unreachable or erroring guardrail refuses the request rather than
# letting it through unscreened. Set `false` only where a guardrail outage must not
# become a gateway outage — failing open is quiet, so alert on it if you do.
stop_on_error: true

# Leave off. Alice answers a redaction as a block on this route, because the
# plugin's response contract carries a verdict and a message but no replacement
# text — see README.md, "Masking".
allow_masking: false

request:
url: https://api.alice.io/v2/evaluate/kong
auth:
location: header
name: af-api-key
value: ${{ env "DECK_ALICE_API_KEY" }}
# The request body Alice's adapter reads; nothing else is consulted.
body:
source: $(source)
content: $(content)
# The application whose policies apply, read off the authenticated
# consumer — never from anything the caller writes, so a developer cannot
# point their own traffic at the application with the loosest policies.
# Set each consumer's Custom ID to the Application ID in Alice (below).
app_id: $(kong.client.get_consumer() and kong.client.get_consumer().custom_id)
# The request whole. With it, Alice reads the conversation the way its
# LiteLLM integration does — the newest turn and its tool calls screened,
# the system prompt and tool definitions recorded — instead of screening
# Kong's one flattened string. Optional: without it, `content` is screened.
request_body: $(kong.request.get_raw_body())
# Optional. Group a conversation or attribute a user from a header. Claude
# Code needs neither: both are read from its own request metadata.
# session_id: $(kong.request.get_header("x-session-id"))
# user_id: $(kong.request.get_header("x-user-id"))

response:
# Both are single field reads on purpose: Alice's verdict is flat so that this
# config needs no Lua function of its own.
block: $(resp.blocked)
block_message: $(resp.message)

metrics:
# A string, on purpose: `categories` is an array and metrics take a string.
block_reason: $(resp.message)
block_detail: $(resp.correlation_id)

# One consumer per application. Its Custom ID is the Application ID in Alice — the field on the
# add-application form, described there as "your own ID for this application".
consumers:
- username: payments-bot
custom_id: payments-bot
keyauth_credentials:
- key: ${{ env "DECK_PAYMENTS_BOT_KEY" }}

verdict Alice returns to Kong

{
"blocked": true,
"message": "Blocked by content policy: hate",
"categories": ["hate"],
"correlation_id": "…"
}

Good to know​

A policy configured to MASK is answered as a block on this route, because Kong's plugin contract carries no replacement text; keep allow_masking false. stop_on_error: true refuses the caller when Alice is unreachable. On a block Kong returns HTTP 400 with the policy message.