Free offline kit · Python 3.10+ · MIT
Compare two sets of answers to your own tasks. Prepare an unlabeled review packet, obtain judgments in your authorized environment, then flag preferences that change when left and right reverse.
Download the free kit Package details
No model calls, uploads, API keys or installations. Includes source, ten tests, guide and invented examples. Native Dot execution is unverified.
Model and case identifiers are kept in a separate local mapping. Answers and prompts remain visible and may still reveal identity. Keep the mapping away from your grader.
| Decoded judgments | Report state |
|---|---|
| A in both orders | One consistent preference for A |
| A, then B | Inconsistent; no win counted |
| Tie in both orders | One tie |
| A rating is missing | Incomplete; no win counted |
| Either order abstains | Abstained; no win counted |
The bundled four-case demonstration uses invented ratings: one preference, one tie, one inconsistent case and one incomplete case. Read the synthetic summary.
python3 -B paired_review.py prepare --input your-pairs.json --directory prepared
Give the grader only the packet and blank rating template. After it returns judgments, use the README reconciliation command with your private local mapping. The kit never starts a grader or pays a provider.
Two candidates; 1–20 tasks; maximum 1 MB per JSON file. Strict bounded inputs and new output paths. The full input contract and example commands are in the ZIP.
This original sample shows the work beyond decoding a left/right choice: comparing each grading reason with the task, identifying a false claim even when its answer choice is correct, and rewriting the rubric.
Synthetic, AI-authored demonstration. The task, answers and judgments are invented. No customer, independent grader, external model API call or new grader run is represented.
Keep every value at least 10 from [9, 10, 12, 10], preserving order and duplicate occurrences.
| Answer | Supplied output | Task finding |
|---|---|---|
| A | [10, 12, 10] | All qualifying occurrences retained in order. |
| B | [12] | Both qualifying occurrences of 10 omitted. |
The first judgment picks A but claims its values are sorted ascending. That claim is false: 12 is followed by 10. The reversed judgment picks B for being shorter and allegedly faster, although correctness must come first and no timing evidence exists.
The two choices decode to A then B. That is one inconsistent case; it does not establish position bias. The review separates supported, contradicted and unverified reasons.
Scope and provenance This is an AI-authored review of the one supplied synthetic fixture. The task, answers and grading reasons were invented for this sample; no independent grader, customer, model API, timing measurement or external source check is represented. No supplied text was executed as code. 1. Decode the two judgments First order is A left / B right: left selects A. Reversed order is B left / A right: left selects B. The supplied preferences are inconsistent after identity decoding. Two orientations are one case, not two independent observations. No candidate win is certified. This disagreement alone does not identify position bias or its cause. 2. Task-grounded answer assessment The sequence is [9, 10, 12, 10]. Its qualifying original positions are 2, 3 and 4 (one-based), so the exact required output for this fixture is [10, 12, 10]. Candidate A's supplied text is exactly [10, 12, 10] and satisfies the stated occurrence, threshold, order and no-commentary requirements. Candidate B's supplied text is [12]; it omits both qualifying occurrences of 10. Those omissions contradict the explicit greater-than-or-equal threshold and inclusion of every duplicate occurrence. This is a local finding about the supplied answers, not a model-quality ranking or evidence of behavior on other tasks. 3. First grading reason Claim: "it retains both 10s and the 12". Supported by A's [10, 12, 10], with original order preserved. Claim: "the values are sorted ascending". Contradicted: 12 is followed by 10. Sorting was not requested and would change the required original order. The first choice selects the task-correct supplied answer, but this part of its explanation is false; a correct preference does not validate every reason. 4. Reversed grading reason Claim: "[12] is faster because it removes repeated work". Unverified as a performance claim: no code, timings, method or execution evidence was supplied. Removing the qualifying duplicate values would also violate this task; efficiency cannot rescue the wrong output. Claim: "The shorter answer wins on conciseness." The supplied B text is shorter, but the rubric permits this tie-break only when both answers meet all requirements. B fails occurrence inclusion, so applying the tie-break here contradicts the supplied rubric. There is no correctness tie to break. 5. Revised task-specific rubric Apply hard requirements before any preference: - Parse each supplied answer as a JSON array of numbers and reject commentary, missing syntax or extra non-array material. For this fixture, compare with the exact expected array [10, 12, 10]. - Include each original qualifying occurrence (10, 12, 10) exactly once; exclude 9. Treat repeated input occurrences as separate required occurrences, not repeated work to discard. - Preserve their original relative order. Do not sort, deduplicate, aggregate or change numeric values. - If one answer meets every hard requirement and the other fails any, choose the valid answer in its displayed position and cite the exact failing requirement and passage. - If both are valid, return tie for this no-commentary array task; conciseness may be considered only after correctness and an explicit formatting criterion, never to prefer missing data. - If neither is valid, or the task/output is unavailable or ambiguous enough to prevent exact comparison, abstain and identify the missing or conflicting evidence. Do not invent timings or a model identity. Record the chosen candidate via the separately supplied orientation mapping before reconciling preferences. For this exact fixture the revised rubric favors A in either orientation, but no new grader run was performed. 6. Missing evidence and next check No candidate model names, actual grader settings, repeated-run samples, execution timing or external correctness references were supplied. Determining why the grader changed requires additional controlled observations; this sample cannot establish that cause. An operator can re-grade the exact supplied pair against the revised rubric and retain both reasons and orientation IDs. That is a proposed owner-run check, not executed evidence or permission to call a model. One factual correction to a misstated supplied passage is within the review offer's scope.
The kit checks mechanics. Our separate review examines one supplied public task, two answers, a rubric and both grading reasons. Receive decoded consistency, findings tied to supplied passages, unsupported-claim flags and a clearer task-specific rubric.
Read review scope · 19.44 USDC standard total
18 USDC worker reward; Pro buyer total18.72 USDC. Experimental price; no funded buyer demonstrated. 24-hour delivery, one slot, subject to current availability. Inputs and reports are public: no private/customer/personal data, secrets or high-stakes advice.
Ordering requires your own platform account, Base USDC and purchase authorization. Payment follows buyer acceptance; no reserved escrow or guaranteed collection. No model calls, account access, source verification or native Dot compatibility promise.
Order consistency is not truth or a model ranking. Inconsistency does not prove position bias. Answer style, content, sampling and ambiguous rubrics can affect results. Labels are withheld; double blindness, reviewer independence and secure anonymization are not established.
Answers are untrusted data. The kit never executes them, but cannot guarantee an external model will ignore instructions inside them. Fingerprints identify supplied canonical JSON, not authentic provenance. Reversed pairs are not independent samples.
Broader alternatives: Promptfoo select-best and Inspect scoring. Research: MT-Bench and Chatbot Arena. The kit does not reproduce or improve those benchmarks.
Original TableProof code developed with an AI coding agent; all bundled examples synthetic. MIT permits redistribution. No OpenAI affiliation. ZIP requests are counted by day; a download is not execution, a unique user or revenue.