← RELSA Desk / API
Tokens

Drive RELSA Desk from your own code

A reading of the computed scores, not of the animals. The model reads the RELSA scores and severity zones your browser (or your script) computed; it never recomputes a number. A score is a fraction of the reference set's maximum deviation, comparable only within that reference frame. KDE zones are candidate, model-specific cut-points, not severity grades under EU Directive 2010/63/EU, and nothing here predicts an animal's course or replaces your humane endpoint criteria.

Everything the web page does is available over HTTP. The page reads a welfare cohort table (one row per animal per time point: an id, a time column, optional labels such as treatment or condition, and readout columns such as body weight, temperature, a clinical score or a biomarker) and computes RELSA scores with its own relsa.js and relsakit.js, following the skill's relsa_score.py and kde_thresholds.py: each variable as a percent of the animal's own baseline, ordinal scores mapped onto the percent scale, a reference model from the group assumed to carry the greatest burden, a weight per variable per time point, the root-mean-square RELSA score, and severity zones from the minima of a kernel density estimate with a bandwidth sweep. Send those facts and get back a verdict (sound, caveated, unreliable) and either a reading of every metric, the animals, the zones and the reporting checklist, or a Python script that reproduces every number with the skill's own functions and runs follow-up checks. The metered run sees only the browser's analysis, never your raw table. RELSA Desk is derived from the agent skill @k-dense-ai/relsa-severity-assessment (k-dense-ai/scientific-agent-skills, K-Dense Inc.).

Two lanes: the task field

taskwhat you getextra input
interpretA reading of each metric (M1..M7, in order: scored rows, the reference set, the highest score, animals above the reference maximum, the KDE thresholds, bandwidth stability, animals whose set of measured variables changes), the animals the scores single out, the severity zones in words, the eight-item reporting checklist, a methods draft, your claims judged against the facts, and what the scores cannot show.none
scriptThe fixes (re-scoring without a variable or with its direction flipped, rebuilding the reference set from a named group, scoring only variables measured at every time point, repeating the KDE across the sweep's bandwidths, excluding baseline rows from the KDE, forecasting one animal) and one complete Python script: SKILL_SCRIPTS = "scripts" on sys.path, COHORT_PATH = "cohort.csv" read with read_relsa_table, the skill's prepare, build_reference, relsa_scores and find_thresholds, an EXPECTED dict of every browser value checked with math.isclose, then the fixes, all inside main().decision: the text of an earlier interpret run (optional)

Both lanes return the same envelope: lane, verdict, headline, tldr, the lane body, next_steps and prescan_responses. Worked requests: interpret, script. The reply shape: output contract.

Input fields

Every field is a string.

fieldrequiredmeaning
taskyesinterpret (a reading of what the scores and zones support: metric readings, animals, zones, an 8-item reporting checklist, a methods draft, claims, cautions) or script (a Python script that reproduces the scores with the skill's own functions, plus fixes, assumptions and checks). These are the only two lanes.
factsyesA JSON-encoded string holding the browser's analysis - see below. The page builds it with RelsaKit.facts in relsakit.js; an API caller must build the same object. The easiest path is to open the app, load your data and copy the run input (or reuse the example below).
titlenoA label for the analysis, up to 160 characters.
contextnoYour notes: the model or procedure, the humane endpoint criteria applied, why this reference set is assumed to carry the greatest burden, what you want to conclude. The checklist can only mark endpoint_criteria and reference_set as stated from here.
questionnoAnswered in tldr as a bullet starting "Answer:".
decisionscript onlyPlain text of an earlier interpret result to act on (the page builds it with Recon.decisionText: "Verdict: ...", the headline, one line per metric reading, animal and zone string, then the next steps).
retry_notenoSet only when re-asking after a malformed reply.

The facts string

facts is a JSON string, not an object: the browser scores the cohort, serialises the result with JSON.stringify and sends that text. Its keys are settings (rows, animals, variables_scored, turned, normalized, score_scales, baseline_time, baseline_rule, reference, dropped, rounding, the kde settings, forecasting, software); reference_model (per variable: turned, max_reached_pct, max_delta_pct, the denominator of every weight); animals (highest peak first: max_relsa, time_of_max, last_relsa, weights_at_max, zone_at_max and more); groups; kde (n, bandwidth, thresholds, modes, zones, sweep, or {"error"}); metrics (M1..); flags (F1.. with severity high / medium / low, category, message, refs); browser_verdict; pipeline (the exact settings in the skill's terms, for the script); expected (the values a reproduction must match: n_scored, maxdelta__<VAR>, relsa__<ID>__<TIME>, kde_n, kde_bandwidth, kde_threshold_count, kde_threshold_<k>), expected_rows (which id and time each relsa__ key names), expected_count; and clipped (what was left out for length).

The skill's own synthetic cohort. The object below is what the page computes for the 6-mouse example cohort that ships with the relsa-severity-assessment skill (not real animals), with the endpoint animals as the reference set and the density taken over the treated animals. The full string is about 8.6 KB; here the animals list and expected_rows are cut after their first items, and the entries whose key is "…" mark the cuts. Send the full string from the page, never this abbreviation:

{
  "settings": {
    "rows": 54,
    "animals": 6,
    "variables_scored": ["weight", "temp", "score", "il6"],
    "turned": ["il6", "score"],
    "normalized": ["weight", "temp", "il6"],
    "score_scales": [
      {"column": "score", "max_score": 8, "baseline_score": 0}
    ],
    "baseline_time": -1,
    "baseline_rule": "the listed baseline time point",
    "reference": {
      "label": "cohort.csv [condition=endpoint]",
      "loaded_from_file": false,
      "groups": ["condition=endpoint"],
      "n_animals": 2,
      "n_rows": 18
    },
    "dropped": [],
    "rounding": "2 decimals on deltas, weights and scores (R package convention)",
    "kde": {
      "filter": ["treatment=treated"],
      "excluded_times": null,
      "bandwidth_rule": "bw.nrd0 (Silverman, R)",
      "n_thresholds_kept": 2,
      "min_zone_fraction": 0.02,
      "grid": 512,
      "cut": 3
    },
    "forecasting": "not run in the browser (ARIMA foRcast is available in the skill's forecast_relsa.py)",
    "software": "relsa-severity-assessment skill scripts, version 1.1 (numpy/pandas/scipy)"
  },
  "reference_model": [
    {"variable": "weight", "turned": false, "max_reached_pct": 82.39968216, "max_delta_pct": 17.60031784},
    {"variable": "temp", "turned": false, "max_reached_pct": 92.78600269, "max_delta_pct": 7.213997308},
    {"variable": "score", "turned": true, "max_reached_pct": 187.5, "max_delta_pct": 87.5},
    {"variable": "il6", "turned": true, "max_reached_pct": 797.7207977, "max_delta_pct": 697.7207977}
  ],
  "animals": [
    {"id": "M01", "in_reference": true, "n_timepoints": 9, "n_scored": 7, "treatment": "treated", "condition": "endpoint", "max_relsa": 1, "time_of_max": 5, "last_relsa": 1, "last_time": 5, "at_peak_at_last_time": true, "largest_weight_at_max": {"variable": "weight", "weight": 1}, "weights_at_max": {"weight": 1, "temp": 1, "score": 1, "il6": 1}, "timepoints_above_1": 0, "zone_at_max": "danger", "zone_at_last": "danger"},
    {"…": "5 more animals (M02, M03, M04, S02, S01), highest peak first - the full string comes from the page"}
  ],
  "groups": [
    {"group": "treatment=treated, condition=endpoint", "animals": 2, "scored_rows": 15, "highest_max_relsa": 1, "highest_animal": "M01", "animals_above_1": 0},
    {"group": "treatment=treated, condition=survivor", "animals": 2, "scored_rows": 18, "highest_max_relsa": 0.53, "highest_animal": "M03", "animals_above_1": 0},
    {"group": "treatment=sham, condition=sham", "animals": 2, "scored_rows": 18, "highest_max_relsa": 0.13, "highest_animal": "S02", "animals_above_1": 0}
  ],
  "kde": {
    "n": 33,
    "bandwidth": 0.150177,
    "bw_nrd0": 0.150177,
    "thresholds": [0.7028],
    "modes": [0.2638, 0.8664],
    "zones": [
      {"zone": "normal", "low": 0, "high": 0.7028, "n": 25, "fraction": 0.758},
      {"zone": "danger", "low": 0.7028, "high": null, "n": 8, "fraction": 0.242}
    ],
    "sweep": [
      {"factor": 0.7, "bandwidth": 0.1051, "thresholds": [0.189, 0.671]},
      {"factor": 0.8, "bandwidth": 0.1201, "thresholds": [0.195, 0.68]},
      {"factor": 0.9, "bandwidth": 0.1352, "thresholds": [0.193, 0.69]},
      {"factor": 1, "bandwidth": 0.1502, "thresholds": [0.703]},
      {"factor": 1.1, "bandwidth": 0.1652, "thresholds": [0.728]},
      {"factor": 1.25, "bandwidth": 0.1877, "thresholds": []}
    ]
  },
  "metrics": [
    {"metric": "scored_rows", "value": 51, "fraction": 0.9444, "basis": "rows with a RELSA score (at least one variable measured), of 54 rows from 6 animals", "id": "M1"},
    {"metric": "reference_set", "value": 2, "basis": "animals in the reference set (condition=endpoint), 18 rows", "id": "M2"},
    {"metric": "highest_relsa", "value": 1, "animal": "M01", "time": "5", "basis": "the highest score of any animal and when it occurred; 1 means the reference set's maximum deviation", "id": "M3"},
    {"metric": "animals_above_reference_max", "value": 0, "animals": [], "basis": "animals whose score exceeded 1 at some time point, i.e. deviated further than anything in the reference set", "id": "M4"},
    {"metric": "kde_thresholds", "value": 1, "thresholds": [0.7028], "bandwidth": 0.150177, "n": 33, "basis": "density minima of the 33 scores in the KDE (Gaussian kernel, bandwidth bw.nrd0)", "id": "M5"},
    {"metric": "bandwidth_stability", "value": 2, "of": 6, "basis": "bandwidths in the sweep (0.7, 0.8, 0.9, 1, 1.1, 1.25 x bw.nrd0) that give the same number of thresholds as the one used", "id": "M6"},
    {"metric": "composition_changes", "value": 0, "animals": [], "basis": "animals whose set of measured variables changes along the trajectory, which moves the score by itself", "id": "M7"}
  ],
  "flags": [
    {"id": "F1", "severity": "low", "category": "small_reference", "message": "The reference set has 2 animal(s). Every weight divides by that set's single most extreme value, so one animal's worst day fixes the scale.", "refs": []},
    {"id": "F2", "severity": "low", "category": "score_mapping", "message": "Score mapping applied: score healthy 0.0 -> 100%, worst 8.0 -> 200%. This is a modelling choice about what a score point is worth against a percent of body weight, and it belongs in the methods.", "refs": ["score"]},
    {"id": "F3", "severity": "medium", "category": "bandwidth_sensitive", "message": "The number of thresholds changes within 10% of the bandwidth used (the sweep keeps it at 2 of 6 bandwidths). Report the sweep, not a bare threshold.", "refs": []},
    {"id": "F4", "severity": "low", "category": "few_scores", "message": "The density rests on 33 scores (the published sepsis analysis used 239). Minima from a sample this small are fragile.", "refs": []},
    {"id": "F5", "severity": "low", "category": "baseline_scores_in_kde", "message": "5 score(s) in the density are exactly 0, usually the baseline rows, which are 0 by construction. The published analysis excluded the baseline time point; consider excluding it here.", "refs": []}
  ],
  "browser_verdict": "caveated",
  "pipeline": {
    "cohort_file": "cohort.csv",
    "id_col": "id",
    "time_col": null,
    "variables": ["weight", "temp", "score", "il6"],
    "turned": ["il6", "score"],
    "normalize": ["weight", "temp", "il6"],
    "score_scale": [
      {"column": "score", "max_score": 8, "baseline_score": 0}
    ],
    "baseline_time": -1,
    "reference": {"filter": [["condition", "endpoint"]], "n_animals": 2, "n_rows": 18},
    "drop": [],
    "round_digits": 2,
    "kde": {
      "filter": [["treatment", "treated"]],
      "exclude_times": null,
      "bandwidth": null,
      "n_thresholds": 2,
      "min_zone_fraction": 0.02,
      "within_data": true
    }
  },
  "expected": {
    "n_scored": 51,
    "maxdelta__weight": 17.60031784,
    "maxdelta__temp": 7.213997308,
    "maxdelta__score": 87.5,
    "maxdelta__il6": 697.7207977,
    "relsa__M01__5": 1,
    "relsa__M02__5": 0.94,
    "relsa__M02__6": 0.94,
    "relsa__M03__3": 0.53,
    "relsa__M03__7": 0.11,
    "relsa__M04__4": 0.41,
    "relsa__M04__7": 0.09,
    "relsa__S02__3": 0.13,
    "relsa__S02__7": 0.04,
    "relsa__S01__1": 0.11,
    "relsa__S01__7": 0.02,
    "kde_n": 33,
    "kde_bandwidth": 0.1501770151,
    "kde_threshold_count": 1,
    "kde_threshold_1": 0.7027551545
  },
  "expected_rows": [
    {"key": "relsa__M01__5", "id": "M01", "time": 5},
    {"key": "relsa__M02__5", "id": "M02", "time": 5},
    {"…": "9 more rows"}
  ],
  "expected_count": 20,
  "clipped": []
}

The free browser page computes this full object for any cohort you paste (CSV or tab-separated, up to the page's row limit). To copy it without writing code, run the page's saved example (it replays for free) or your own analysis and press Download .json: the file carries the exact facts object under browser. Send it back as a string: json.dumps(facts), JSON.stringify(facts) or your language's equivalent. Keep the keys and values the browser produced: the reply is reconciled against them, and the script lane copies expected into its reproduction check.

Building the body

The simplest way to get a body that matches the page byte for byte is to run the page's own module in Node. relsakit.js loads relsa.js (the scoring and KDE engine) from the same folder, and both export themselves with module.exports. The input is the page's Save set .json file: the table under data plus the settings the form holds (idCol, timeCol, variables, turned, normalize, scoreScale as COL=MAX[:BASELINE], baselineTime, referenceGroup as col=value, referenceJson, drop, precision, kdeGroup, kdeBaseline include / exclude, nThresholds, bandwidth, minZone) and the title, context and question.

// make-body.js - build the exact body the page sends, with the page's own code.
// Save https://relsa-desk.skillsafe.ai/relsakit.js and relsa.js next to this file, then:
//   node make-body.js my-set.json interpret > body.json
//   node make-body.js my-set.json script "Verdict: caveated. ..." > body.json
// my-set.json is the page's "Save set .json" download (table + settings + notes).
const fs = require("fs");
const K = require("./relsakit.js");            // requires ./relsa.js itself
const [setFile, lane = "interpret", decision = ""] = process.argv.slice(2);
const set = JSON.parse(fs.readFileSync(setFile, "utf8"));
set.lane = lane;
if (lane === "script") set.decision = decision;
const A = K.analyze(set);
if (A.empty) throw new Error(A.errors.join("; ") || "no rows to score");
const body = K.mustBeObject(K.buildInput(A, set));
fs.writeFileSync("cohort.csv", K.cohortCsv(A));  // the table the script lane reads
console.error("browser verdict:", A.hint, "| animals:", A.animals.length, "| scored rows:", A.scored,
  "| flags:", A.flags.map(f => f.id + " " + f.category).join(", "));
console.error("idempotency key: relsa-desk:" + body.task + ":" + K.hashInput(body) + ":a1");
fs.writeSync(1, JSON.stringify(body));
# Or build the body in any language from a facts object you already hold, for example the
# "browser" key of the page's "Download .json" export. facts must go in as a STRING.
import json

export = json.load(open("relsa-desk-interpret.json"))   # the page's .json download
facts = export["browser"]
body = {
    "task": "interpret",
    "title": "Example cohort: endpoint animals as reference",
    "context": "The synthetic 6-mouse cohort that ships with the skill (not real animals). ...",
    "question": "Which animals came closest to the reference maximum?",
    "facts": json.dumps(facts, separators=(",", ":"), ensure_ascii=False),
}
json.dump(body, open("body.json", "w"), ensure_ascii=False)

Base URL and the envelope

Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses the same envelope, so one helper covers the whole API:

{"ok": true, "data": {"job_id": "job_...", "status": "queued"}}
{"ok": false, "error": {"code": "payment_required", "message": "..."}}

The token is minted for this app (the guest endpoint takes {"slug":"relsa-desk"} in its body), so no slug header is needed afterwards. Send it as Authorization: Bearer ….

The input object IS the request body. There is no {"input": …} wrapper. A wrapped body is answered with an unknown field 'input' warning, and the model never sees your facts.

Error codes

statuscodewhat to do
400validation_errorA field is missing or the wrong type. Every field is a string: facts must be a JSON-encoded string, not an object.
401unauthorizedThe token is missing, malformed or expired. Get a new one from the token page.
402payment_requiredThe balance is below min_credits. Call /estimate first and top up.
403forbiddenThe token is valid but not for this app, or a guest token tried a metered run. A guest cannot run; sign in for a personal token.
404not_foundUnknown job id, or the app slug does not exist.
409conflictThe same Idempotency-Key was replayed with a different body. Change the key or send the original input.
429rate_limitedToo many requests. Back off and retry; do not tight-loop.
5xxinternalA server-side failure. Retry with the SAME Idempotency-Key so you are not billed twice.

1. A tiny client

One helper that sends the token, unwraps data and raises on ok: false. The token comes from the token page (Copy token or Copy shell export); step 2 covers the kinds of token and minting one from code.

# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="relsa-desk"
TOKEN="$SKILLSAFE_TOKEN"   # from https://relsa-desk.skillsafe.ai/tokens.html

call() {                  # call <path> [json-body]
  if [ -n "$2" ]; then
    curl -sS -X POST "$BASE/$1" \
      -H "Authorization: Bearer $TOKEN" \
      -H "Content-Type: application/json" \
      -d "$2"
  else
    curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
  fi
}

2. Get a token

The easiest route is the token page: it shows the token this browser already holds, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. A guest token, minted with POST /guest and {"slug":"relsa-desk"}, can call /me and /estimate; the run is metered, so /run and /run-stream need a personal token.

# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
#   https://relsa-desk.skillsafe.ai/tokens.html
#   export SKILLSAFE_TOKEN="..."
#
# To mint a guest token from the command line instead. A guest token is enough
# for /me and /estimate; a run needs a personal token from signing in.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
  -H "Content-Type: application/json" -d '{"slug":"relsa-desk"}'
# {"ok":true,"data":{"token":"…","subject_type":"guest"}}

3. Check the session and the balance

call me
# {"ok":true,"data":{"subject_type":"user","username":"you","credits":51234}}

4. Price the run (free)

/estimate returns the model binding and the credits a run would reserve. It creates no job and charges nothing. Expect model_alias gpt-terra and markup_bps 1000 (a 10% markup). hold_credits is a reservation, not the price: it is held against your balance while the run executes and released afterwards. min_credits is the least balance that can start a run. What you actually pay is charged_credits, reported on the finished job and in the done event, and it is usually far lower than the hold. The body is the input object itself, with no {"input": …} wrapper. /estimate does not validate the body, so check the shape yourself: an object whose every value is a string, task equal to interpret or script, facts non-empty, and facts a JSON string that parses to an object (this is what the page's own guard, RelsaKit.mustBeObject, refuses to spend without).

# body.json is the input object itself - no {"input": ...} wrapper. Build it with
# make-body.js above, or by hand. estimate does not validate it, so check the shape first:
python3 -c 'import json;b=json.load(open("body.json"));assert isinstance(b,dict) and b.get("task") in ("interpret","script") and all(isinstance(v,str) for v in b.values()) and all(b.get(k,"").strip() for k in ("facts",)) and isinstance(json.loads(b["facts"]),dict)'
INPUT=$(cat body.json)

call estimate "$INPUT"
# {"ok":true,"data":{"model":"...","model_alias":"gpt-terra",
#   "markup_bps":1000,"hold_credits":...,"min_credits":...,"sponsor_enabled":false,
#   "warnings":[]}}
#
# estimate creates no job and charges nothing. hold_credits is RESERVED, not the
# price; charged_credits after the run is the actual cost, usually far lower.

5. Run it, then poll

POST /run returns a job_id; poll GET /jobs/{id} until it is terminal. The reply is a string at data.output.output: JSON.parse it (step 7). Send an Idempotency-Key built from the lane, a hash of the input and the attempt number, relsa-desk:<lane>:<hash>:a<attempt> (for example relsa-desk:interpret:1a5xi21gged1p:a1), so a retried request returns the same job instead of billing a second run. Use one key per distinct input: a changed table, setting, reference set or note (so changed facts) or a changed reading is a new hash, the same cohort in the other lane is a new key, and replaying an old key with a different body is a 409. The page uses RelsaKit.hashInput(body) for the hash (it covers task, title, context, facts, decision and question; make-body.js prints the key); any stable digest of the body works from other languages. Leave retry_note out of the hash and bump the attempt instead.

# Always send an Idempotency-Key derived from the input. A retried request with
# the same key returns the SAME job instead of billing a second run.
LANE=$(printf '%s' "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["task"])')   # interpret or script
KEY="relsa-desk:$LANE:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"

JOB=$(curl -sS -X POST "$BASE/run" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')

while :; do
  OUT=$(call "jobs/$JOB")
  STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
  [ "$STATUS" = "succeeded" ] && break
  [ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
  sleep 2
done

# {"ok":true,"data":{"job_id":"job_...","status":"succeeded",
#   "output":{"output":"{\"lane\":\"interpret\",\"verdict\":\"caveated\",\"headline\":\"...\", ...}"},
#   "charged_credits":...,"truncated":false}}
printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])' > reply.json

6. Or stream it

POST /run-stream takes the same body and headers and answers with server-sent events: job (the job id), delta (chunks of the reply) and done (the status, charged_credits, truncated and, when present, the full output). A browser page may receive only tick heartbeats and then done, never a delta, so take the reply from done.output.output when it is there, fall back to the concatenated deltas, and fall back again to GET /jobs/{id}.

# Server-sent events. `delta` events carry chunks of the reply; `done` carries the
# status, charged_credits and the truncated flag. Ignore `tick` heartbeats.
curl -N -X POST "$BASE/run-stream" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -H "Accept: text/event-stream" \
  -d "$INPUT"

# event: job    {"job_id":"job_..."}
# event: delta  {"text":"{\"lane\":\"interpret\",\"verdict\":\"caveated\",\"headline\":\"The"}
# event: done   {"status":"succeeded","charged_credits":...,"truncated":false}

7. Parse the reply

The reply is a JSON object serialised as a string. Parse it, then check the lane.

# The reply is a JSON string inside data.output.output. Pull it out and parse it:
printf '%s' "$JOB" | python3 -c 'import sys,json;r=json.loads(json.load(sys.stdin)["output"]["output"]);print(r["verdict"],r["headline"])'

Invariants worth asserting

The output contract

The model returns one JSON object as the job's output text. Every key of the lane's contract is present; empty sections are [].

{
  "lane": "interpret" | "script",
  "verdict": "sound" | "caveated" | "unreliable",
  "headline": "one sentence",
  "tldr": ["2-5 bullets; one starts \"Answer:\" when a question was asked"],
  // interpret:
  "metrics": [{"id": "M1", "reading": "..."}],          // one per facts.metrics item, same order
  "animals": ["1-6 strings: the highest peaks, animals still at their peak, groups compared"],
  "zones": ["1-4 strings: thresholds, bandwidth, number of scores, what the sweep shows"],
  "checklist": [{"item": "directionality", "status": "stated|partly|missing|not_applicable", "note": "..."}],
                                                        // exactly 8, in the order listed above
  "methods": "one paragraph, with [AUTHOR_INPUT_NEEDED: ...] where an item is missing",
  "claims": [{"claim": "...", "support": "supported|partly|not_supported", "why": "..."}],
  "cautions": ["1-4 strings on what the scores cannot show"],
  // script:
  "fixes": [{"fix": "...", "why": "...", "refs": "F1"}],  // 1-6; refs "" for a plain step
  "script": "import math\nimport sys\n...",               // one complete Python script, under 9000 characters
  "assumptions": ["1-4 strings"],
  "checks": ["1-4 strings"],
  // both:
  "next_steps": ["1-5 concrete actions"],
  "prescan_responses": [{"ref": "F1", "verdict": "confirmed|dismissed", "note": "..."}]
}

The script lane's script follows a fixed order: set SKILL_SCRIPTS = "scripts", check it is a folder and put it on sys.path; set COHORT_PATH = "cohort.csv", check it is an existing local file and read it with read_relsa_table(COHORT_PATH, id_col=..., time_col=...) from pipeline.id_col and pipeline.time_col; map each pipeline.score_scale column with score_to_percent; call prepare(frame, normalize=..., baseline_time=...); build the reference from the pipeline.reference.filter rows with build_reference (or ReferenceModel.from_json on reference.json when the reference was loaded from a file); score with relsa_scores(prepared, reference, drop=..., round_digits=...); run find_thresholds on the finite relsa values the pipeline.kde settings select; define EXPECTED and ROWS (each expected_rows key to its (id, time)); compare every key with math.isclose(got, want, rel_tol=1e-6, abs_tol=1e-9), printing any mismatch; then apply the fixes, printing what changes.

Worked example: interpret

The skill's synthetic 6-mouse cohort from the facts section above (not real animals), exactly as the page sends it for its first example. The browser scores 51 of 54 rows, takes the two endpoint animals (M01, M02) as the reference set, finds one density threshold at 0.7028 over 33 scores and raises F3 (medium) because only 2 of the 6 bandwidths in the sweep keep the same number of thresholds, so its read is caveated. The body, with facts abbreviated (send the full string from make-body.js or the page):

{
 "task": "interpret",
 "title": "Example cohort: endpoint animals as reference",
 "context": "The synthetic 6-mouse cohort that ships with the relsa-severity-assessment skill (not real animals). Temperature and body weight fall under burden; IL-6 rises; the 0-8 clinical score is mapped with a score scale. The reference set is the two endpoint-reaching animals, assumed to carry the greatest burden in this model.",
 "question": "Which animals came closest to the reference maximum, and can we use the density threshold as a danger line?",
 "facts": "{\"settings\":{\"rows\":54,\"animals\":6,\"variables_scored\":[\"weight\",\"temp\",\"score\",\"il6\"],\"turned\":[\"il6\",\"score\"]},\"…\":\"the rest of settings, reference_model, animals, groups, kde, metrics, flags, pipeline and expected - the full string comes from the page\",\"browser_verdict\":\"caveated\",\"expected_count\":20,\"clipped\":[]}"
}

An abbreviated reply (the job's output.output, parsed; "..." marks cuts):

{
  "lane": "interpret",
  "verdict": "caveated",
  "headline": "Against the two endpoint animals as reference, M01 reached the reference maximum and M02 came close, but the single density threshold at 0.703 is bandwidth-sensitive.",
  "tldr": [
    "51 of 54 rows are scored; the reference set is the 2 endpoint animals (condition=endpoint).",
    "No animal exceeded 1, the reference set's maximum deviation.",
    "Answer: M01 (1 at time 5) and M02 (0.94 at time 5) came closest; the threshold at 0.703 holds for only 2 of 6 bandwidths, so it is a candidate cut-point to report with its sweep, not a danger line."
  ],
  "metrics": [
    {"id": "M1", "reading": "51 of 54 rows (94.44%) have a RELSA score."},
    {"id": "M2", "reading": "The reference set is 2 animals with condition=endpoint, 18 rows."},
    {"id": "M3", "reading": "M01 has the highest score, 1 at time 5: it defines the reference maximum."},
    {"id": "M4", "reading": "0 animals went above 1, so none deviated further than the reference set."},
    {"id": "M5", "reading": "The density of 33 scores has 1 threshold, 0.703, at bandwidth 0.150 (bw.nrd0)."},
    {"id": "M6", "reading": "Only 2 of 6 bandwidths in the sweep give the same number of thresholds."},
    {"id": "M7", "reading": "0 animals change their set of measured variables along the trajectory."}
  ],
  "animals": [
    "M01 peaked at 1 at time 5 and M02 at 0.94 at time 5; both are in the reference set and were at their peak at their last time point.",
    "Among the survivors M03 peaked at 0.53 at time 3; the sham group's highest peak was 0.13 (S02).",
    "..."
  ],
  "zones": [
    "One threshold at 0.703 splits the 33 treated-animal scores into normal (25) and danger (8) at bandwidth 0.150.",
    "The sweep gives two thresholds at 0.7, 0.8 and 0.9 x bw.nrd0 and none at 1.25 x, so the threshold is bandwidth-sensitive."
  ],
  "checklist": [
    {"item": "directionality", "status": "stated", "note": "il6 and score are turned; the notes say temperature and weight fall and IL-6 rises."},
    {"item": "baseline", "status": "stated", "note": "The listed baseline time point, -1."},
    {"item": "score_mapping", "status": "stated", "note": "score 0-8 mapped onto the percent scale (max_score 8, baseline_score 0)."},
    {"item": "reference_set", "status": "stated", "note": "The 2 endpoint animals, assumed in the notes to carry the greatest burden."},
    {"item": "endpoint_criteria", "status": "missing", "note": "The humane endpoint criteria applied are not given."},
    {"item": "forecasting", "status": "not_applicable", "note": "No forecast is reported."},
    {"item": "thresholds", "status": "partly", "note": "Bandwidth bw.nrd0 and 33 scores are stated; the threshold changes across the sweep."},
    {"item": "software", "status": "stated", "note": "relsa-severity-assessment skill scripts, version 1.1."}
  ],
  "methods": "Severity was assessed with RELSA using body weight, temperature and IL-6 as a percent of each animal's baseline at time -1 and a 0-8 clinical score mapped onto the percent scale. ... Humane endpoints were [AUTHOR_INPUT_NEEDED: the humane endpoint criteria applied]. ...",
  "claims": [
    {"claim": "The density threshold can serve as a danger line.", "support": "not_supported", "why": "The number of thresholds changes across the sweep (M6: 2 of 6) and the density rests on 33 scores; a KDE zone is a candidate cut-point, not a decision rule."}
  ],
  "cautions": [
    "The scores are relative to these 2 endpoint animals and cannot be compared with another study's.",
    "The zones are not severity categories under EU Directive 2010/63/EU."
  ],
  "next_steps": ["State the humane endpoint criteria applied.", "Report the bandwidth sweep with the threshold.", "..."],
  "prescan_responses": [
    {"ref": "F1", "verdict": "confirmed", "note": "The reference set has 2 animals."},
    {"ref": "F2", "verdict": "confirmed", "note": "The score mapping is a modelling choice for the methods."},
    {"ref": "F3", "verdict": "confirmed", "note": "Only 2 of 6 bandwidths keep the same number of thresholds."},
    {"ref": "F4", "verdict": "confirmed", "note": "The density rests on 33 scores."},
    {"ref": "F5", "verdict": "confirmed", "note": "5 scores in the density are exactly 0."}
  ]
}

A reply must answer F1 to F5 once each in prescan_responses, read M1 to M7 in order, give the 8 checklist items in order, judge the danger-line claim in question against kde, and stay at caveated or tighter unless it dismisses F3.

Worked example: script

The page's second example, the 22-mouse synthetic CLP-like cohort (generated for the page, not real animals: telemetry temperature and activity, body weight and a 0-8 clinical score, 8 endpoint animals as the reference set, the baseline time point excluded from the density), sent to the script lane with a short decision (the page fills it with Recon.decisionText of the earlier interpret reply; it may be empty). The browser's read is sound, with two low flags: F1 (score mapping) and F2 (one score within a grid step of the threshold). No question is sent in this lane. The body, with facts abbreviated:

{
 "task": "script",
 "title": "Synthetic CLP-like cohort, telemetry + weight + score",
 "context": "Synthetic cohort generated for this example (not real animals): a CLP-like sepsis model with telemetry temperature and activity, body weight and a 0-8 clinical score, 8 endpoint animals, 8 survivors and 6 shams. Endpoint animals were removed at the humane endpoint, so their later rows are empty. The reference set is the endpoint group, assumed to carry the greatest burden.",
 "decision": "Verdict: sound.\nAgainst the 8 endpoint animals as reference, E05 peaked highest at 0.97 at time 6 and one density threshold at 0.759 holds for 4 of 6 bandwidths.\n- M1: 181 of 198 rows are scored.\n- ...\nNext steps:\n- Report the bandwidth sweep with the threshold.",
 "facts": "{\"settings\":{\"rows\":198,\"animals\":22,\"variables_scored\":[\"temp\",\"act\",\"weight\",\"score\"],\"turned\":[\"score\"],\"normalized\":[\"temp\",\"act\",\"weight\"]},\"…\":\"the rest of settings, reference_model, animals, groups, kde, metrics and flags - the full string comes from the page\",\"browser_verdict\":\"sound\",\"pipeline\":{\"cohort_file\":\"cohort.csv\",\"id_col\":\"id\",\"time_col\":null,\"variables\":[\"temp\",\"act\",\"weight\",\"score\"],\"turned\":[\"score\"],\"normalize\":[\"temp\",\"act\",\"weight\"],\"score_scale\":[{\"column\":\"score\",\"max_score\":8,\"baseline_score\":0}],\"baseline_time\":-1,\"reference\":{\"filter\":[[\"condition\",\"endpoint\"]],\"n_animals\":8,\"n_rows\":72},\"drop\":[],\"round_digits\":2,\"kde\":{\"filter\":[],\"exclude_times\":[-1],\"bandwidth\":null,\"n_thresholds\":null,\"min_zone_fraction\":0.02,\"within_data\":true}},\"expected\":{\"n_scored\":181,\"maxdelta__temp\":9.887760556,\"maxdelta__act\":87.70301624,\"maxdelta__weight\":17.64950166,\"maxdelta__score\":100,\"relsa__E05__6\":0.97,\"relsa__E04__5\":0.96,\"relsa__E02__5\":0.94,\"relsa__E07__5\":0.94,\"relsa__E01__6\":0.93,\"relsa__E03__5\":0.93,\"kde_n\":159,\"kde_bandwidth\":0.09298578192,\"kde_threshold_count\":1,\"kde_threshold_1\":0.7585894213},\"expected_rows\":[{\"key\":\"relsa__E05__6\",\"id\":\"E05\",\"time\":6},{\"key\":\"relsa__E04__5\",\"id\":\"E04\",\"time\":5},{\"…\":\"4 more rows\"}],\"expected_count\":15,\"clipped\":[]}"
}

An abbreviated reply; the script string is shown decoded below it:

{
  "lane": "script",
  "verdict": "sound",
  "headline": "The script rebuilds the endpoint reference from cohort.csv with the skill's functions, checks all 15 expected values, then repeats the density across the sweep and re-scores without score.",
  "tldr": ["The reproduction check covers n_scored, 4 maxdelta values, 6 peak scores and 4 KDE values.", "..."],
  "fixes": [
    {"fix": "Repeat the KDE across the bandwidth factors of facts.kde.sweep and print the thresholds.", "why": "One score lies within a grid step of the threshold at 0.759, so its zone depends on the estimate.", "refs": "F2"},
    {"fix": "Re-score without score and print each animal's peak score before and after.", "why": "The score mapping is a modelling choice; this shows how much the peaks rest on it.", "refs": "F1"}
  ],
  "script": "import math\nimport sys\nfrom pathlib import Path\n...",
  "assumptions": ["cohort.csv is the file the page wrote for this cohort.", "The skill's scripts folder sits at ./scripts."],
  "checks": ["The reproduction prints no MISMATCH line.", "How the peaks move without score, and whether the threshold near 0.759 persists across the sweep."],
  "next_steps": ["Save the page's cohort.csv next to relsa_check.py and run python relsa_check.py.", "..."],
  "prescan_responses": [
    {"ref": "F1", "verdict": "confirmed", "note": "Answered by the second fix."},
    {"ref": "F2", "verdict": "confirmed", "note": "Answered by the first fix."}
  ]
}
# relsa_check.py - reproduce the browser's RELSA scores and zones (abbreviated reply script)
import math
import sys
from pathlib import Path

import numpy as np

SKILL_SCRIPTS = "scripts"
if not Path(SKILL_SCRIPTS).is_dir():
    raise SystemExit(f"{SKILL_SCRIPTS}/ not found: point SKILL_SCRIPTS at the skill's scripts folder")
sys.path.insert(0, SKILL_SCRIPTS)

from _common import read_relsa_table, score_to_percent  # noqa: E402
from kde_thresholds import find_thresholds  # noqa: E402
from relsa_score import build_reference, prepare, relsa_scores  # noqa: E402

COHORT_PATH = "cohort.csv"
VARIABLES = ["temp", "act", "weight", "score"]
TURNED = ["score"]
NORMALIZE = ["temp", "act", "weight"]
BASELINE_TIME = -1
KDE_EXCLUDE_TIMES = [-1]
SWEEP = [0.7, 0.8, 0.9, 1, 1.1, 1.25]

EXPECTED = {
    "n_scored": 181,
    "maxdelta__temp": 9.887760556,
    "maxdelta__act": 87.70301624,
    "maxdelta__weight": 17.64950166,
    "maxdelta__score": 100,
    "relsa__E05__6": 0.97,
    "relsa__E04__5": 0.96,
    "relsa__E02__5": 0.94,
    "relsa__E07__5": 0.94,
    "relsa__E01__6": 0.93,
    "relsa__E03__5": 0.93,
    "kde_n": 159,
    "kde_bandwidth": 0.09298578192,
    "kde_threshold_count": 1,
    "kde_threshold_1": 0.7585894213,
}
ROWS = {
    "relsa__E05__6": ("E05", 6),
    "relsa__E04__5": ("E04", 5),
    "relsa__E02__5": ("E02", 5),
    "relsa__E07__5": ("E07", 5),
    "relsa__E01__6": ("E01", 6),
    "relsa__E03__5": ("E03", 5),
}


def kde_values(scores):
    rows = scores[~scores["time"].isin(KDE_EXCLUDE_TIMES)]
    values = rows["relsa"].to_numpy(dtype=float)
    return values[np.isfinite(values)]


def main():
    if not Path(COHORT_PATH).is_file():
        raise SystemExit(f"{COHORT_PATH} not found: save it from the RELSA Desk page first")
    frame = read_relsa_table(COHORT_PATH, id_col="id", time_col=None)
    frame["score"] = score_to_percent(frame["score"], max_score=8, baseline_score=0)
    prepared = prepare(frame, normalize=NORMALIZE, baseline_time=BASELINE_TIME)
    ref_rows = prepared[prepared["condition"].astype(str) == "endpoint"]
    reference = build_reference(ref_rows, variables=VARIABLES, turned=TURNED, baseline_time=BASELINE_TIME)
    scores = relsa_scores(prepared, reference, drop=[], round_digits=2)
    values = kde_values(scores)
    kde = find_thresholds(values, bandwidth=None, n_thresholds=None, min_zone_fraction=0.02)

    got = {"n_scored": int(np.isfinite(scores["relsa"].to_numpy(dtype=float)).sum())}
    for var in VARIABLES:
        got[f"maxdelta__{var}"] = reference.maxdelta[var]
    for key, (animal, time) in ROWS.items():
        row = scores[(scores["id"].astype(str) == animal) & (scores["time"] == time)]
        got[key] = float(row["relsa"].iloc[0])
    got["kde_n"] = kde.n
    got["kde_bandwidth"] = kde.bandwidth
    got["kde_threshold_count"] = len(kde.thresholds)
    for k, t in enumerate(kde.thresholds, start=1):
        got[f"kde_threshold_{k}"] = t
    bad = [k for k, want in EXPECTED.items()
           if k not in got or not math.isclose(got[k], want, rel_tol=1e-6, abs_tol=1e-9)]
    for k in bad:
        print(f"MISMATCH {k}: got {got.get(k)!r}, expected {EXPECTED[k]!r}")
    print("reproduction:", "OK" if not bad else f"{len(bad)} mismatch(es)")

    # Fix 1 (F2): repeat the KDE across the sweep's bandwidth factors.
    for factor in SWEEP:
        r = find_thresholds(values, bandwidth=kde.bandwidth * factor, min_zone_fraction=0.02)
        print(f"bandwidth x{factor}: thresholds {[round(t, 3) for t in r.thresholds]}")
    # Fix 2 (F1): re-score without score and compare each animal's peak.
    without = relsa_scores(prepared, reference, drop=["score"], round_digits=2)
    before = scores.groupby("id")["relsa"].max()
    after = without.groupby("id")["relsa"].max()
    for animal in before.index:
        print(f"{animal}: peak {before[animal]} -> {after[animal]} without score")


if __name__ == "__main__":
    main()

The reply's script must put all 15 keys of facts.expected into EXPECTED with the browser's values, check each with math.isclose, and answer F1 and F2 with fixes from the allowed list. Save the page's cohort.csv next to the script, point SKILL_SCRIPTS at the skill's scripts folder and run it with python relsa_check.py.

Truncation and partial results

If your balance sits between min_credits and hold_credits, the run still executes with a smaller output cap and the job carries "truncated": true. The JSON may then stop mid-object: close it (the page's Recon.closeJson does this) and show the sections that arrived, saying how many of the lane's sections were recovered, rather than treating a clipped reply as complete. A clipped script is not runnable; re-run instead.