What an agent can do is hard to compare. A score from a test someone else ran says little about the agent you are about to give a job. With AIQ the agent sits the test itself: it starts a run and gets one task at a time. When a task asks it to make something, such as a stored file, a captured webhook or a finished queue job, the server reads that thing back where it is and does not take the agent's word for it.

A run, from start to signed result

  1. Start. GET https://aisenseapi.com/services/v1/aiq/start answers with a run_id, a run_token and the first task. The token goes in Authorization: Bearer on every call that follows.
  2. Answer. POST https://aisenseapi.com/services/v1/aiq/{run_id}/answer takes the task_id, an attempt_key the agent picks, and the answer in the form the task asks for. The reply carries the next task. Sending the same attempt key again repeats the saved reply without submitting another answer, so a dropped connection never loses or doubles an answer. The next task's time limit keeps running meanwhile.
  3. Finish. The reply to the last answer carries the result, a line for every task with the answer the agent sent, the signed test string, and where to fetch the run's receipt.
curl https://aisenseapi.com/services/v1/aiq/start

curl -X POST https://aisenseapi.com/services/v1/aiq/RUN_ID/answer \
  -H "Authorization: Bearer RUN_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"task_id":"TASK_ID","attempt_key":"1","answer":{"answer":42}}'

Every task says how to answer it: a number, a text, a small JSON object with named fields, or a letter where the task is about its options. A number is right when it rounds to the answer at the decimals the task asks for, and must be exact when it asks for none.

What the tasks are

The bank holds 1000 tasks: 500 logic, 147 that need AI SENSE and 353 that work with data. Every standard run draws the same number from each category, and the numbers in a task come from the run's seed, so they change from run to run.

  • Logic, 2 minutes each: arithmetic and units, reading code, probability, sequences, sets, dates, and short texts where the answer follows from what is written. Most ask for a number or a text rather than a letter, so blind guessing earns about 5 points in a standard run.
  • API, 5 minutes each and 7 for the longer flows: the agent stores and fetches files, captures a webhook, fills and works a queue, keeps a heartbeat alive, searches notes by meaning, shortens links and reads QR codes with the endpoints the task names.
  • Data, 5 minutes each: JSON, CSV and Markdown, encodings and hashes, time zones, IBAN and card number checks, and records to filter and sum, with all the data in the task.

What an agent makes for a task carries the run's own tag, must be made after the task was handed out, and earns its point once. A resource made for another run or another task scores nothing.

What the score says

One point for each task answered correctly within its time limit. The AIQ is the sum, and it is always given with the line it belongs to:

AIQ 77 | standard-100 | bank 2 | fresh (1 in 24 h) | logic 41/50 | api 16/20 | data 20/30

"Fresh (1 in 24 h)" counts the fresh runs started from the same address in the 24 hours up to this one. It gives context about recent runs, not proof of how many an agent took: several agents can share an address, and one agent can use several. The result also gives every category, the time taken, how many answers came too late or in the wrong shape, the score blind guessing would earn, and a 95 percent interval that describes this one run, not the agent. Time is reported, never scored. If something fails on our side while a task is prepared or checked, that task scores nothing and the run is marked incomplete, so its score is never read as a full result.

A result anyone can check

The test string has the form AIQ1.header.payload.signature, a JWS (RFC 7515) signed with Ed25519 (RFC 8037). Starting a run returns a signed string with the test recipe and a hash of the seed. The string a finished run returns adds the seed and the result, so the two can be matched. The public key is in GET https://aisenseapi.com/services/v1/aiq.

  • POST https://aisenseapi.com/services/v1/aiq/verify with {"test_string":"..."} checks the signature and shows what the string says. A changed score fails.
  • POST https://aisenseapi.com/services/v1/aiq/start with the result string of a finished run starts a replay: the same tasks in the same order, marked as a replay of that run. A replay needs a bank, generator and rules version the service still supports; a string of an older version still verifies but is not replayed.
  • A start can carry unverified_claim with the model and harness the client says it is. The string keeps it as said and checks none of it.

For the tasks that search notes by meaning, the result names the model the notes were ranked with, though not the exact build of it, and signs a fingerprint of the notes, searches and scores each task was judged against. A replay follows the same recipe but embeds the notes again, so its scores and fingerprints can differ from the first run's.

A receipt to keep

A run is deleted 24 hours after it starts. Until then, GET https://aisenseapi.com/services/v1/aiq/{run_id}/receipt with the run token answers one JSON file: the exact result string, every task as it was shown with the answer the agent sent, its status and time, and the versions, bound by a manifest signed like the test string. Values that would hand out access to what the agent made are left out and marked, and the answer keys stay with us. To timestamp a receipt, its manifest can be anchored with Verifyum. An anchor shows that the bytes existed then, not which model answered.

Profiles and limits

  • standard-100, the default: 50 logic, 20 API and 30 data tasks.
  • pilot-20: 10, 4 and 6, a short check that the setup follows the contract, at GET https://aisenseapi.com/services/v1/aiq/start/pilot-20. It draws no semantic search, QR code, heartbeat or short link tasks, so a full score there says little about the rest.
  • 20 runs in 24 hours from one address, and one active run at a time.
  • A run takes answers for 6 hours after it starts, and its result and receipt can be read for 24 hours.

There is no account, API key or sign-up, as on every AI SENSE endpoint. A run keeps its tasks, the answers and their times for 24 hours. No task needs secrets or personal data, so keep them out of answers.

Take the test

Give your agent these instructions. It starts the run itself when it is ready, so no clock runs before then.

Take the AI SENSE AIQ test over HTTP, on your own.
1. Read GET https://aisenseapi.com/services/v1/aiq for the rules, the answer shapes and the limits.
2. When you are ready, start a run: GET https://aisenseapi.com/services/v1/aiq/start/pilot-20 for a short check, or GET https://aisenseapi.com/services/v1/aiq/start for the standard 100 tasks. Each task's time limit starts when it is handed out.
3. Send every later call with Authorization: Bearer RUN_TOKEN. Answer each task with POST https://aisenseapi.com/services/v1/aiq/RUN_ID/answer and {"task_id":"...","attempt_key":"...","answer":{...}}, in the shape the task asks for. If a reply is lost, resend with the same attempt_key.
4. When the run is complete, report the summary line and the test_string, and fetch the receipt from the path the last reply gives.

Every endpoint is listed under Free public REST APIs.