A decision model does one narrow thing. You give it a situation and a fixed set of questions, and it gives back an answer to each with probabilities a program can act on: yes or no, one of several options, or a place on a scale. It writes no text to parse. That is the same job DecideAgent Optimal has done with rules since September, so when Cloudflare published two such models in the open, we made room for them.

What Cloudflare released

During its Birthday Week Cloudflare released Clef and Clef-flash, open source under the Apache 2.0 licence and published on Hugging Face. Clef is built on a 27-billion-parameter Qwen model and Clef-flash on a 9-billion one. Both answer three kinds of question: noul, a yes or no given as a number from 0 to 1, choice and score, and both follow the API of Jev, the model that defined the category.

By Cloudflare's own measurements Clef-flash answers in 38.8 ms at the median and 122.4 ms at the 95th percentile, against 524.1 ms for Jev, and the larger Clef in 209.3 ms. We have not measured those numbers ourselves, and ours will differ, because our copy runs on our own machine.

Two ways to ask Decide

RulesClef
You sendFacts, and weighted rules for each answerThe situation in words, the question, and what each answer means
Questionsyes_no, choice, scalenoul, choice, score
You get backProbabilities from the weights, how clear-cut they are, a next step of act, review or hold for the answer given, and the rules that heldThe model's answer, probabilities and its own confidence, with no next step
RunsInside the APIOn our own inference machine, one call at a time
TodayOpen to everyoneSwitched off until measured

The endpoint is the same: POST https://aisenseapi.com/services/v1/decide. Leave model out, or send "rules", and Decide answers from rules exactly as before. Send "model": "clef" and the questions carry instructions and criteria in plain words instead of rules:

{
  "model": "clef",
  "state": { "message": "The production API returns HTTP 500 and blocks checkout." },
  "questions": {
    "urgent": {
      "type": "noul",
      "instructions": "Does this need urgent attention?",
      "criteria": {
        "true": "An active production failure blocks business.",
        "false": "The issue can wait."
      }
    }
  }
}

There is no fallback from one to the other. An unknown model is refused with a 400, and the API checks every field of the model's reply before it passes any of it on.

Why the model is switched off for now

Our Clef runs on a single inference machine, not on Workers AI, and that machine does one inference at a time. We have not measured it yet, so until we have, a public model call answers 503 with a fix that says rules still work. The operator measures it first with two fixed probes, a one-question health check and a four-question support ticket, and sets the limits from what they show. The starting limits are deliberately small:

  • 8 KiB in a request, at most 4 questions and 8 options or levels in each
  • 2 model calls a minute and 20 a day from one address, and 300 a day in all
  • one call in flight, no queue, at least 10 seconds between starts, and 30 seconds before the API gives up

A model call sends the situation and the questions to that machine, so keep credentials and sensitive personal data out of them. The API keeps neither the request bodies nor the answers. As on every endpoint, the access log records each request with the address, and the daily request limit counts requests per address. Model calls also leave quota counters keyed by a hash of the address, and timings, as the privacy page says.

Try it without writing JSON

Try Decide is a page for anyone who wants to see what Decide does before writing code. Pick an example, such as whether to refund a customer or bring an umbrella, or fill in your own facts and rules. The JSON request is written for you as you go, and you can edit it. Press Decide and the answer comes back in plain words, with the rules that held, and as JSON. The Clef tab works the same way and is ready for the day the model opens.

Decide is Agent Optimal

Decide answers with a next step an agent can follow, so it now carries the Agent Optimal mark beside Agent Wake, Agent Queue, Agent Inbox, Heartbeat and Lease. Through the MCP server, the decide tool takes the same optional model.

Make a decision

Start in the browser, or read the reference for rules, conditions and the model questions.