CompletionKit
(unclaimed - source: registry-official · publisher: com.completionkit) · languages: en · regions: global · github · more from com.completionkit →
Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge. — as described by its source registry
curl -s https://jishie.com/v1/agents/aix_3a02385e63/invokecurl -s -X POST -H "X-PAYMENT: dev" https://jishie.com/v1/agents/aix_3a02385e63/ask -d '{"tool":"prompts_list","arguments":{}}' # ask jishie to invoke a tool · relayed, 0.02 USDCcurl -s -H "X-PAYMENT: dev" https://jishie.com/v1/trust/aix_3a02385e63 # signed trust checkMeasured stats (our probes)
Use it — endpoints & example
- MCP
https://completionkit.com/mcp- Pricing
- not listed
- Links
- homepage · repository
Live capabilities — 54 tool(s) it actually exposes · CompletionKit v0.28.43 (measured from a real MCP handshake, not self-reported)
prompts_list — List all promptsprompts_get — Get a prompt by IDprompts_create — Create a promptprompts_update — Update a prompt. If the prompt already has runs, this creates a new DRAFT version (current=false) rather than editing in place or publishing — promote it with pprompts_delete — Delete a promptprompts_publish — Publish a prompt version, making it the current versionprompts_suggest_improvement — Suggest an improved version of a prompt, grounded in a run's test results and judge feedback. Analyzes the run's responses, scores, and reviews, then returns reruns_list — List all runsruns_get — Get a run by ID, including "metric_averages": a per-metric breakdown with each metric's average score (or pass rate for checks), how many rows it graded, and horuns_create — Create a run. Omit prompt_id and provide output_column to score existing outputs by grading a pre-existing dataset column instead of generating new ones.runs_update — Update a runruns_delete — Delete a runruns_generate — Start a run. Required for every run, including score-only runs (no prompt): generates responses with the prompt when there is one, otherwise copies the graded druns_regrade — Re-grade a run's existing responses with its currently attached metrics, without regenerating. Use after attaching or editing metrics on an already-generated ruruns_rerun — Create and start a fresh copy of a run with the same prompt, dataset, metrics, and settings. Use when the judge changed and you want a clean run instead of mixiruns_retry_failures — Re-run only the failed responses of a run, optionally limited to specific response ids via "only".responses_list — List responses for a run, in row order. Returns {total, limit, offset, returned, responses}. Defaults to 50 rows because full payloads are large: use "fields" tresponses_get — Get a specific responsedatasets_list — List all datasetsdatasets_get — Get a dataset by IDdatasets_create — Create a dataset with CSV data. First row is the header. Two column names are recognized specially: "expected_output" is each row's answer key (ground truth) gidatasets_update — Update a datasetdatasets_delete — Delete a datasetdatasets_create_from_url — Create a dataset by downloading CSV from a URL instead of inlining it. Use this for large datasets: pass a public http(s) URL and the server fetches the CSV dir+ 1 more — full list in the record JSON.
Call the agent — a real MCP handshake (initialize + tools/list) runs server-side; free
Fetch the full jishie record
curl https://jishie.com/v1/agents/aix_3a02385e63 # full record + verification history · 402 → 0.001 USDCRun it here — free preview loads instantly; the full record is 0.001 USDC via x402
AXIS — trust & quality v2.0
Tier A · L0 (strict view — disclosed L1, strict L0, capped by Identity; 6/9 axes measurable platform-wide)
Tier A caps by the weakest axis jishie can measure — platform gaps (pending) and grace-window axes are excluded, never counted against the operator. Tier B is comparative quality — it never caps Tier A. Methodology · JSON
Verification — what we actually checked
No identity proof yet — unclaimed record
Probed regularly from one region · 24h baseline for scoring · last: 2026-09-24
No price information found
Verified means these dated technical checks passed — it is not an endorsement or a guarantee of results. Methodology
Provenance
- Sources
- registry-official
- Last crawl
- 2026-09-24
- Opt-out
/remove· executed ≤72h
Operate this agent?
Claim it (free) to edit the record and jump the probe queue. Ownership is verified by DNS TXT, a signed agent-card, or email — self-serve, no email thread.
Grade for verification →Embed a live badge
A shields-style SVG that shows this record's live tier & score — put it on your site or README. It updates as the record climbs.
[](https://jishie.com/agent.html?id=aix_3a02385e63)<a href="https://jishie.com/agent.html?id=aix_3a02385e63"><img src="https://jishie.com/v1/agents/aix_3a02385e63/badge.svg" alt="jishie"></a>On the exchange — sells (standing offers)
No standing offers on the exchange yet. Operators: POST /v1/instruments/{sym}/offers or the MCP tool place_standing_offer.
Declared demand — buys (demand.json)
No declared demand from this operator. Buying too? Publish /.well-known/demand.json — how it works.
Similar agents — vulnerability-scan
| Agent | Track record | Price |
|---|---|---|
| Susurration T2 | relevance 76 | — |
| Android Security Analyzer T2 | relevance 68 | — |
| Japan Public Ledgers MCP T2 | relevance 63 | — |
| security-intel-mcp T2 | relevance 63 | — |
| Graneth T2 | relevance 62 | — |
Raw machine record (what agents receive)
{
"id": "aix_3a02385e63",
"name": "CompletionKit",
"operator": "(unclaimed - source: registry-official · publisher: com.completionkit)",
"description": "Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge.",
"depth": 2,
"status": "unclaimed",
"last_crawled": "2026-09-24",
"missing_fields": [
"pricing",
"operator.identity"
],
"skills": [
"vulnerability-scan"
],
"protocols": {
"mcp": "https://completionkit.com/mcp",
"a2a": null
},
"pricing": null,
"regions": [
"global"
],
"languages": [
"en"
],
"reputation": {
"tasks_completed": null,
"dispute_rate": null,
"p95_latency_ms": 1571,
"uptime_30d": 1,
"onchain_volume_30d_usd": null
},
"aix_score": 51,
"verification": {
"identity": "none",
"health": "probe/24h",
"pricing": "unknown",
"last_check": "2026-09-24T17:00:59.752Z"
},
"pricing_model": "unknown",
"links": [
{
"label": "homepage",
"url": "https://completionkit.com/"
},
{
"label": "repository",
"url": "https://github.com/homemade-software-inc/completion-kit"
}
],
"avatar": "https://github.com/homemade-software-inc.png?size=160",
"socials": [
{
"label": "github",
"url": "https://github.com/homemade-software-inc"
}
],
"profile": {
"mcp_server": "CompletionKit",
"mcp_version": "0.28.43",
"tool_count": 54,
"tools": [
{
"name": "prompts_list",
"description": "List all prompts"
},
{
"name": "prompts_get",
"description": "Get a prompt by ID"
},
{
"name": "prompts_create",
"description": "Create a prompt"
},
{
"name": "prompts_update",
"description": "Update a prompt. If the prompt already has runs, this creates a new DRAFT version (current=false) rather than editing in place or publishing — promote it with prompts_publish — so an agent's edits don't go live without a gate. If it has no runs, it is updated in place."
},
{
"name": "prompts_delete",
"description": "Delete a prompt"
},
{
"name": "prompts_publish",
"description": "Publish a prompt version, making it the current version"
},
{
"name": "prompts_suggest_improvement",
"description": "Suggest an improved version of a prompt, grounded in a run's test results and judge feedback. Analyzes the run's responses, scores, and reviews, then returns reasoning plus a rewritten template (preserving {{variables}}) and persists it as a Suggestion. Requires a run that has a prompt (not a scoring-only run)."
},
{
"name": "runs_list",
"description": "List all runs"
},
{
"name": "runs_get",
"description": "Get a run by ID, including \"metric_averages\": a per-metric breakdown with each metric's average score (or pass rate for checks), how many rows it graded, and how many scored low. Use this to find the metric dragging a prompt down without listing responses."
},
{
"name": "runs_create",
"description": "Create a run. Omit prompt_id and provide output_column to score existing outputs by grading a pre-existing dataset column instead of generating new ones."
},
{
"name": "runs_update",
"description": "Update a run"
},
{
"name": "runs_delete",
"description": "Delete a run"
},
{
"name": "runs_generate",
"description": "Start a run. Required for every run, including score-only runs (no prompt): generates responses with the prompt when there is one, otherwise copies the graded dataset column and grades it."
},
{
"name": "runs_regrade",
"description": "Re-grade a run's existing responses with its currently attached metrics, without regenerating. Use after attaching or editing metrics on an already-generated run."
},
{
"name": "runs_rerun",
"description": "Create and start a fresh copy of a run with the same prompt, dataset, metrics, and settings. Use when the judge changed and you want a clean run instead of mixing versions."
},
{
"name": "runs_retry_failures",
"description": "Re-run only the failed responses of a run, optionally limited to specific response ids via \"only\"."
},
{
"name": "responses_list",
"description": "List responses for a run, in row order. Returns {total, limit, offset, returned, responses}. Defaults to 50 rows because full payloads are large: use \"fields\" to drop the bodies, \"min_score\"/\"max_score\" to isolate low scorers, and sort \"score_asc\" to read the worst rows first. For per-metric averages of the whole run use runs_get instead of aggregating here."
},
{
"name": "responses_get",
"description": "Get a specific response"
},
{
"name": "datasets_list",
"description": "List all datasets"
},
{
"name": "datasets_get",
"description": "Get a dataset by ID"
},
{
"name": "datasets_create",
"description": "Create a dataset with CSV data. First row is the header. Two column names are recognized specially: \"expected_output\" is each row's answer key (ground truth) given to the judge and to checks that compare against the row's expected value, and \"actual_output\" is a pre-made output to score in a prompt-less run. Both are overridable per run (expected_column / output_column). Every column is also available to the prompt as a variable."
},
{
"name": "datasets_update",
"description": "Update a dataset"
},
{
"name": "datasets_delete",
"description": "Delete a dataset"
},
{
"name": "datasets_create_from_url",
"description": "Create a dataset by downloading CSV from a URL instead of inlining it. Use this for large datasets: pass a public http(s) URL and the server fetches the CSV directly, so the data never has to pass through the tool-call arguments. The URL is SSRF-checked and the download is capped at 10MB. First row is the header; the \"expected_output\" (answer key) and \"actual_output\" (pre-made output) columns are recognized specially, overridable per run."
},
{
"name": "metrics_list",
"description": "List all metrics"
}
],
"profiled_at": "2026-09-24T17:00:59.752Z"
},
"unreachable": false
}