(i) Audio-grounded initialization. Rubrics are generated from the raw waveform, so every criterion is anchored in
the acoustic evidence actually present in the clip.
(ii) Evolving rubrics. At each step, saturated criteria are pruned and harder ones are elicited from the rollouts,
so the evaluation standard co-evolves with the policy.
(iii) Overthinking penalty. A linear length penalty counterbalances the rubric reward, keeping the reasoning
informative but concise.
$R_i = R^{\text{out}}_i + \gamma\, R^{\text{rub}}_i + \delta\, R^{\text{over}}_i$