Skip to content

Fine-tuning and post-training data

Rubrics as rewards: expert rubric data for RL in non-verifiable domains

Quick answer

Rubric-based rewards replace a single pass/fail verifier with a set of prompt-specific, weighted criteria that a judge model scores and combines into one reward [1]. They let RL reach tasks where no exact answer exists, such as clinical advice, legal memos or support resolutions. The bottleneck is the rubric itself: expert-written criteria are expensive and hard to scale [2], so buyers need source material that encodes how professionals actually grade work, plus scored examples to calibrate the judge.

By SourceX Editorial · Updated

How rubric rewards extend verifiable-reward RL

Rubric rewards generalize RLVR by swapping an exact-match or unit-test verifier for a checklist of criteria graded per response. In RLVR, as used in Tulu 3, a prompt ships with a reference answer and a deterministic checker [4]; the reward is binary and cheap to compute. The Rubrics as Rewards paper frames rubric-guided RL as the multi-criteria, prompt-specific extension of that setup and reports gains on HealthBench and GPQA-Diamond over plain LLM-as-judge baselines [1].

The practical difference is where the ground truth lives. With RLVR it sits in a reference answer, which our guide to RLVR datasets with prompts, reference answers and verifiers covers in detail. With rubric rewards it sits in criteria such as "states the contraindication with anticoagulants" or "does not promise a refund timeline the policy does not support," each of which a judge model can check independently.

Checklist rewards are the same idea with lighter structure: the checklist is derived from the instruction itself, and each item is either judged by a model or, where possible, checked by code (a word limit, a required JSON key, a citation format). Mixing programmatic checks with judged items keeps the cheap, deterministic part of RLVR wherever it applies and spends judge calls only on criteria that need reading comprehension.

What a reward-grade rubric item contains

A reward-grade rubric item is a self-contained, binary-checkable statement with a weight and a polarity. Synthetic rubrics guided by reference answers of the kind used in health benchmarks such as HealthBench, which the Rubrics as Rewards authors use to report results [1], show the pattern: many short criteria per prompt, each with a point value reflecting its importance, and negative points for behavior to penalize. A judge model then decides, item by item, whether each criterion is met.

Items that work as rewards share five properties:

  • Atomic. One fact or behavior per item, so the judge returns a clean yes or no.
  • Self-contained. The judge should not need hidden context, a lookup or the expert's memory to apply it.
  • Weighted. Points encode importance; a missed red-flag symptom should cost more than a missed courtesy.
  • Signed. Negative items penalize harmful content, hallucinated policy or excessive hedging, which closes the most common reward-hacking routes.
  • Grounded. Each item traces to a source: a guideline, a policy clause, a QA scorecard line or an expert's written rationale.

Items that fail are usually holistic ("response is high quality"), compound ("accurate and concise"), or unverifiable from the response text alone. Those reintroduce the noise of a generic reward model and give the policy room to optimize style instead of substance.

Per-prompt rubrics versus generic rubrics

Per-prompt rubrics give the strongest signal, while generic rubrics are cheaper and transfer across tasks but reward surface features. Rubrics as Rewards builds its rubrics per prompt rather than relying on one universal standard [1]. A generic rubric ("cites a source," "uses plain language") can be reused across thousands of prompts, but it cannot tell the policy which fact was required in this case.

Most production programs mix the two: a small fixed layer of generic criteria for safety, tone and format, and a prompt-specific layer for substance. The prompt-specific layer is where expert time goes and where purchased source data has the most leverage. OpenRubrics responds to the same cost problem by generating rubrics contrastively from preferred and rejected responses and filtering noisy criteria [2], which still requires a human-grounded seed set to validate against.

If your team also builds evaluation rubrics, keep the two sets separate. Our guide to designing evaluation rubrics with domain experts covers held-out evaluation; training rubrics that leak into evaluation make both numbers meaningless.

Aggregating item scores into a scalar reward

Aggregation is an unresolved design choice, and it changes what the policy learns. The open-problem literature contrasts all-or-nothing rewards, which pay only when every criterion is met, with partial-credit schemes that sum weighted items [3]. All-or-nothing is sparse and stalls early training; weighted sums are dense but let a policy farm easy items while skipping the critical one.

Common mitigations include:

  • Gating. Treat a handful of critical items as hard constraints; if any fails, cap the reward regardless of other items.
  • Normalization. Divide achieved points by the maximum positive points achievable for that prompt, so long rubrics do not dominate the batch.
  • Clipping negatives. Floor the reward at zero per prompt, or keep negative totals, depending on whether you want penalties to dominate the gradient.
  • Implicit grading. Give the judge the full rubric and ask for one holistic score; this is cheaper but less auditable than per-item verdicts.

Decide aggregation before you buy or commission rubrics, because it determines what the data must contain: gating needs an explicit "critical" flag, normalization needs complete point values, and per-item grading needs atomic items.

Calibrating the judge against expert judgments

The judge model is part of the reward function, so it must be calibrated against expert labels before RL starts. Do not assume a strong general model agrees with your experts on domain criteria; measure it. A practical loop: collect responses spanning strong, weak and adversarial outputs, have two or three experts mark each criterion met or not met, run the judge on the same items, and compute per-item agreement such as Cohen's kappa.

Disagreement is diagnostic. Low agreement on a specific item usually means the criterion is ambiguous, not that the judge is weak; rewrite it and re-test. Track false "met" verdicts separately from false "not met," since the former are what a policy will exploit. Re-run calibration whenever you change judge model, prompt template or rubric version.

This is where business QA data earns its place. Contact-center and claims QA scorecards already pair weighted criteria with graded real interactions, and the scored examples serve as a ready calibration set. See licensing QA scorecards for AI training and training data for reward models. Annotator credentials matter equally; verifying domain-expert annotators covers qualification tests and ongoing accuracy checks.

Illustrative rubric record for reward training

A usable rubric record bundles the prompt, weighted signed criteria, provenance for each criterion and expert-graded calibration responses. The schema below is a starting point for a data request or an internal spec.

Illustrative example: invented to show structure; it does not describe an available dataset.

{
  "rubric_id": "rbr-claims-00412",
  "rubric_version": "1.2",
  "domain": "insurance_claims_support",
  "prompt": "Customer asks why a water-damage claim was partially denied and what to do next.",
  "context_refs": ["policy_excerpt_HO3_sec_I_perils", "claim_note_redacted"],
  "aggregation": {"method": "weighted_sum_normalized", "gate_on_critical": true},
  "criteria": [
    {"id": "c1", "text": "Identifies gradual seepage as the excluded cause cited in the claim note", "points": 8, "critical": true, "source": "policy_excerpt"},
    {"id": "c2", "text": "Explains the appeal or re-inspection route and the documents required", "points": 6, "critical": false, "source": "qa_scorecard_line_4"},
    {"id": "c3", "text": "States a payout amount or timeline not present in the context", "points": -10, "critical": true, "source": "expert_rationale"},
    {"id": "c4", "text": "Uses plain language without policy jargon left unexplained", "points": 2, "critical": false, "source": "generic_layer"}
  ],
  "calibration_set": [
    {"response_id": "r1", "expert_labels": {"c1": 1, "c2": 1, "c3": 0, "c4": 1}, "raters": 3, "agreement": "unanimous"},
    {"response_id": "r2", "expert_labels": {"c1": 0, "c2": 1, "c3": 1, "c4": 1}, "raters": 3, "agreement": "2_of_3"}
  ],
  "deidentification": {"method": "names_policy_numbers_replaced", "sample_checked": true}
}

Where expert rubric source data comes from

The best raw material for rubrics is evidence of how a business already grades professional work. Useful sources include QA scorecards with graded interactions, supervisor review notes on tickets or case files, audit checklists and their findings, clinical or legal review comments, and style or policy guides that reviewers enforce. Each pairs explicit criteria with real outputs judged against them.

Pair that with a representative prompt set; real-world prompt sets for post-training explains how to keep training prompts close to production traffic. Expert preference pairs are complementary rather than a substitute; domain-expert preference data covers when pairwise judgments beat rubrics. For graded written work in education, see rubric-scored written responses.

Operational records carry personal data. Grading notes and claim files routinely include names, account numbers and free-text details, and health records need HIPAA de-identification by Safe Harbor or Expert Determination [5]; our comparison of Safe Harbor and Expert Determination for AI training explains which to require.

Buyer checklist for rubric reward data

Before you commit, a short diligence pass catches most of the problems that show up later as reward hacking or flat curves.

CheckWhat to ask forFailure it prevents
Item atomicity50 random criteria reviewed for single, checkable claimsNoisy judge verdicts
Weights and signsPoint values on every item, with negatives presentPolicy farming easy items
ProvenanceSource per criterion (policy, scorecard, rationale)Invented or unsupported criteria
Calibration labelsMulti-rater met/not-met labels on sample responsesUncalibrated judge
Rater credentialsQualification and agreement recordsNon-expert criteria
Train/eval separationDisjoint prompt IDs and rubric familiesLeaked evaluation
Rights and de-identificationOwnership, consents, recorded methodUnlicensed or identifiable data
VersioningRubric version and change logIrreproducible reward

The fine-tuning and post-training data hub links related guides on preference, demonstration and verifier data. If the source material you need is held by operating businesses rather than annotation vendors, you can describe the graded records you need to SourceX; requests are sourced case by case and a request does not guarantee a match.

Failure modes to watch during rubric-reward training

Most rubric-reward failures trace back to the data, not the optimizer. Watch for these patterns in reward curves and sampled rollouts:

  • Checklist stuffing. The policy restates every plausible criterion in long, list-like answers. Counter it with negative items for irrelevant content and a length-aware generic layer.
  • Judge sycophancy. The judge marks items met when the response asserts compliance ("I have included the contraindications") without doing it. Calibration sets should include such adversarial responses.
  • Criterion drift. Rubric versions change mid-run, so rewards across checkpoints are not comparable. Pin rubric_version in every rollout log.
  • Coverage gaps. Prompts whose rubrics are thin earn near-maximum reward for weak answers. Flag prompts with few positive points and either enrich or drop them.

Sampling rollouts that score at the top of the distribution and having an expert read them is the cheapest early-warning system available.

Sourcing business data behind expert rubrics

SourceX sources operational datasets from US companies on request, including support and sales histories, documents, and finance and legal workflows, and manages the licensing process. Every dataset is rights-reviewed and delivered under a license defining records, uses, term and delivery, with personal details removed or replaced before delivery; nothing is contracted until a supplier agrees. Describe the rubric source data you need at the SourceX buyer page.

Sources

  1. arXiv (2507.17746), "Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains" (2025). https://arxiv.org/pdf/2507.17746
  2. arXiv, "OpenRubrics: Towards Scalable Synthetic Rubric Generation" (2025). https://arxiv.org/abs/2510.07743
  3. Emergent Mind, "Alternative reward computation methods for rubric-based RL in instruction following". https://www.emergentmind.com/open-problems/alternative-reward-computation-methods-rubric-rl
  4. arXiv (2411.15124), "Tulu 3: Pushing Frontiers in Open Language Model Post-Training" (2024). https://arxiv.org/pdf/2411.15124
  5. U.S. Department of Health and Human Services, Office for Civil Rights, "Guidance Regarding Methods for De-identification of Protected Health Information in Accordance with the HIPAA Privacy Rule" (2012). https://www.hhs.gov/hipaa/for-professionals/special-topics/de-identification

Tell us what your models need

Share scope, volume, language, format, timing and licensing requirements.

Request data