Guardrails for agents
Decide whether an action may run on its own, needs a look, or has to wait for a person.
Logic - Decisions
Send the facts and your rules, and get typed decisions back: yes or no, one of several options, or a level on a scale. Each answer has a probability, a confidence, a recommended action and the rules that produced it. There is no model inside, so every number can be worked out by hand, and nothing is stored.
Send a state, the facts to decide on, and named questions. Each question has rules, and each rule is a condition and a weight. This one routes a support ticket and decides whether the refund can go through at once:
curl -s -X POST https://aisenseapi.com/services/v1/decide \
-H "Content-Type: application/json" \
-d '{
"state": {
"ticket": {
"subject": "Mug arrived broken, I want my money back",
"amount": 30,
"tier": "gold",
"order_age_days": 3,
"has_photo": true
}
},
"questions": {
"team": {
"type": "choice",
"options": {
"returns": [
{"if": {"ticket.subject": {"contains": "broken"}}, "weight": 3},
{"if": {"ticket.subject": {"contains": "return"}}, "weight": 2}
],
"billing": [
{"if": {"ticket.subject": {"contains": "money"}}, "weight": 1},
{"if": {"ticket.subject": {"contains": "charged"}}, "weight": 2}
],
"shipping": [{"if": {"ticket.subject": {"contains": "tracking"}}, "weight": 2}]
}
},
"refund_now": {
"type": "yes_no",
"bias": -2,
"rules": [
{"if": {"ticket.amount": {"lte": 50}}, "weight": 2},
{"if": {"ticket.tier": {"eq": "gold"}}, "weight": 1},
{"if": {"ticket.order_age_days": {"lte": 14}}, "weight": 1},
{"if": {"ticket.has_photo": {"eq": true}}, "weight": 1}
]
}
}
}'{
"answers": {
"team": {
"type": "choice",
"choice": "returns",
"probabilities": {"returns": 0.8438, "billing": 0.1142, "shipping": 0.042},
"confidence": 0.7657,
"action": "review",
"because": [
{"option": "returns", "rule": 0, "weight": 3},
{"option": "billing", "rule": 0, "weight": 1}
]
},
"refund_now": {
"type": "yes_no",
"answer": "yes",
"probability": 0.9526,
"confidence": 0.9051,
"action": "act",
"because": [
{"rule": 0, "weight": 2},
{"rule": 1, "weight": 1},
{"rule": 2, "weight": 1},
{"rule": 3, "weight": 1}
]
}
}
}The refund gets −2 + 2 + 1 + 1 + 1 = 3, and sigmoid(3) = 0.9526. Its confidence, 0.9051, is above 0.9, so the action is act: it can go through without anyone looking. The team gets softmax over 3, 1 and 0, and a confidence of 0.7657 says review. because lists the rules that held, including the word "money" that pulled a little towards billing.
| Type | Answers with | How it is worked out |
|---|---|---|
yes_no | answer, probability | sigmoid(bias + the weights of the rules that held) |
choice | choice, probabilities | softmax over each option: its prior plus the weights of its rules that held |
scale | level, expected, probabilities | a choice over ordered levels; expected counts the first level as 0 |
Every answer also has confidence, action and because. A weight is evidence in log-odds: positive pulls towards the answer, negative away, and 1 multiplies the odds by about 2.7. bias is where a yes_no starts, 0 unless you set it, and priors give options or levels a head start.
A condition is an object of fields and tests, and every field must hold. A field is a dotted path into the state, where a number picks from a list: items.0.sku. any is a list of conditions of which one must hold, and not holds when its condition does not.
| Test | Holds when the value |
|---|---|
eq, ne | equals, or does not equal, the given text, number, true, false or null. 1 equals 1.0; true is not 1 |
gt, gte, lt, lte | is a number above, at least, below or at most the given number |
in | equals one of a list of up to 100 values |
exists | is there (true) or is not (false) |
prefix | is text that starts with the given text, whatever the case |
contains | is text that contains the given text, whatever the case, or a list that holds the given item |
A path that leads nowhere fails every test except exists: false, so ne on a missing field does not hold. There are no regular expressions, so no condition can make the server spend minutes on one match.
Confidence is (n × the largest probability − 1) / (n − 1): 1 when everything is on one answer, 0 when it is spread evenly. For yes_no it is |2p − 1|. The action is act from act_at, 0.9 unless you set it, review from review_at, 0.5 unless you set it, and hold below. Set the thresholds by what a wrong answer costs: a refund of a few euros can act at 0.9, a deletion should not act at all. When two options tie at the top, the first one listed wins and tied names them, so a split is never hidden.
A cheap gate in front of an agent that runs commands. The default is yes, and dangerous patterns pull it down.
curl -s -X POST https://aisenseapi.com/services/v1/decide \
-H "Content-Type: application/json" \
-d '{
"state": {"command": "git push --force origin main", "inside_repo": true, "tests_passed": true},
"questions": {
"auto_run": {
"type": "yes_no",
"bias": 3,
"rules": [
{"if": {"command": {"contains": "--force"}}, "weight": -4},
{"if": {"command": {"contains": "rm -rf"}}, "weight": -6},
{"if": {"command": {"contains": "| sh"}}, "weight": -6},
{"if": {"command": {"contains": " main"}}, "weight": -1},
{"if": {"inside_repo": {"eq": false}}, "weight": -3},
{"if": {"tests_passed": {"eq": true}}, "weight": 1}
]
}
}
}'{
"answers": {
"auto_run": {
"type": "yes_no",
"answer": "no",
"probability": 0.2689,
"confidence": 0.4621,
"action": "hold",
"because": [{"rule": 0, "weight": -4}, {"rule": 3, "weight": -1}, {"rule": 5, "weight": 1}]
}
}
}3 − 4 − 1 + 1 = −1 gives 0.2689: no, with a confidence of 0.46, so hold and ask a person. The same request with git status answers yes at 0.982 and act.
A scale with four levels, acting at 0.8, because paging the person on call is right a little more often than it is wrong.
curl -s -X POST https://aisenseapi.com/services/v1/decide \
-H "Content-Type: application/json" \
-d '{
"state": {"alert": {"service": "payments", "error_rate": 0.12, "customers_affected": 340}},
"questions": {
"severity": {
"type": "scale",
"act_at": 0.8,
"levels": {
"low": [],
"normal": [{"if": {"alert.error_rate": {"lt": 0.05}}, "weight": 2}],
"high": [
{"if": {"alert.error_rate": {"gte": 0.05}}, "weight": 2},
{"if": {"alert.service": {"eq": "payments"}}, "weight": 1}
],
"urgent": [
{"if": {"alert.error_rate": {"gte": 0.1}}, "weight": 2},
{"if": {"alert.customers_affected": {"gte": 100}}, "weight": 2},
{"if": {"alert.service": {"eq": "payments"}}, "weight": 1}
]
}
}
}
}'{
"answers": {
"severity": {
"type": "scale",
"level": "urgent",
"expected": 2.8529,
"probabilities": {"low": 0.0059, "normal": 0.0059, "high": 0.1178, "urgent": 0.8705},
"confidence": 0.8273,
"action": "act",
"because": [
{"level": "high", "rule": 0, "weight": 2},
{"level": "high", "rule": 1, "weight": 1},
{"level": "urgent", "rule": 0, "weight": 2},
{"level": "urgent", "rule": 1, "weight": 2},
{"level": "urgent", "rule": 2, "weight": 1}
]
}
}
}The cheapest layer first: rules, a small model, a large one, a person. Here the large model and the person tie, and the answer says so.
curl -s -X POST https://aisenseapi.com/services/v1/decide \
-H "Content-Type: application/json" \
-d '{
"state": {"text_chars": 1800, "language": "no", "legal_terms": true, "tier": "gold"},
"questions": {
"handler": {
"type": "choice",
"options": {
"rules_only": [{"if": {"text_chars": {"lte": 200}}, "weight": 2}],
"small_model": [
{"if": {"text_chars": {"lte": 2000}}, "weight": 1},
{"if": {"language": {"in": ["en"]}}, "weight": 1}
],
"large_model": [
{"if": {"text_chars": {"gt": 2000}}, "weight": 2},
{"if": {"language": {"in": ["no", "sv", "da"]}}, "weight": 1},
{"if": {"legal_terms": {"eq": true}}, "weight": 1}
],
"human": [
{"if": {"legal_terms": {"eq": true}}, "weight": 1},
{"if": {"tier": {"eq": "gold"}}, "weight": 1}
]
}
}
}
}'{
"answers": {
"handler": {
"type": "choice",
"choice": "large_model",
"tied": ["large_model", "human"],
"probabilities": {"rules_only": 0.0541, "small_model": 0.147, "large_model": 0.3995, "human": 0.3995},
"confidence": 0.1993,
"action": "hold",
"because": [
{"option": "small_model", "rule": 0, "weight": 1},
{"option": "large_model", "rule": 1, "weight": 1},
{"option": "large_model", "rule": 2, "weight": 1},
{"option": "human", "rule": 0, "weight": 1},
{"option": "human", "rule": 1, "weight": 1}
]
}
}
}A confidence of 0.20 gives hold. The rules say plainly that they cannot tell, and the code can then send the job to the safer of the two.
Every refusal has error, naming the field, and fix, saying what to send instead. Nothing is stored whatever the outcome.
| Status | When |
|---|---|
| 400 | A field is missing, of the wrong kind or out of range, or the body is not valid JSON. The error names the place, such as questions.team.options.returns[0].weight |
| 405 | Any method but POST |
| 413 | The body is over 64 KiB, or the rules hold more than 5000 tests |
| 415 | The body is not sent as application/json |
| 429 | The service-wide limit of 5000 requests per IP per 24 hours |
Decide whether an action may run on its own, needs a look, or has to wait for a person.
Send tickets, alerts and jobs to the right place, and see why they went there.
Act only when the evidence is strong, with a threshold that fits what a mistake costs.
Answer the clear cases with rules, and hand only the uncertain ones to a model or a person.
The state and the rules are tested inside the worker and gone when the answer is. Nothing is written, stored or sent anywhere, and the access log line holds the path and nothing from the body. Limits: 64 KiB of JSON, 64 questions, 100 options, 20 levels, 50 rules per list, weights from −100 to 100, 5000 tests and 8 levels of any and not in one request. The service-wide ceiling is 5000 requests per IP per 24 hours.