CriticGen: Generation-Aware Evaluation as Actionable Feedback
Source note: Preprints. Not peer reviewed at time of posting.
arXiv:2609.05439v1 Announce Type: new Abstract: Current evaluation methods for large language models are coarse-grained and decoupled from generation, producing generic explanations that fail to provide actionable feedback for model improvement. We propose CriticGen, a fine-grained, generation-aware evaluation framework that turns evaluation into actionable control for answer improvement. CriticGen first generates sample-specific evaluation dimensions and scoring criteria under high-level categories such as subjective, objective, and self-derived constraints. These criteria then serve as a dynamic rubric for jointly producing a score, a reason, an executable refinement suggestion, and a refined answer. This rubric-conditioned refinement process enables models to diagnose flaws and perform targeted answer improvement. Experimental results show that fine-grained evaluation should be both instance-specific and actionable. CriticGen induces higher-quality rubrics, improving relevance/coverage from 3.33/4.03 to 3.97/4.24. CriticGen also achieves the best score correlations, with 0.9556 Pearson and 0.9560 Spearman, and raises the F1 of criterion-grounded reasons and executable suggestions from 0.6369/0.5994 to 0.7554/0.7900. Crucially, its feedback translates into reliable answer improvement, improving 73.17% of answers with a 93.28% non-degradation rate.
Every model that read this
| Model | Provider | Stage | Score | Conf. | Latency | Prompt | When |
|---|---|---|---|---|---|---|---|
| Llama 3.3 70B | Meta | analysis | +75 | 80% | 3839ms | v1.0.0 / m1.0.1 | 2026-09-10 09:19 |
| GPT-4.1 mini | OpenAI | consensus | +20 | 85% | 3779ms | v1.0.0 / m1.0.1 | 2026-09-11 00:19 |
| Claude Sonnet 5 | Anthropic | consensus | +8 | 70% | 5744ms | v1.0.0 / m1.0.1 | 2026-09-11 00:19 |
CriticGen's evaluation framework improves model performance and answer quality.
CriticGen offers actionable, fine-grained evaluation feedback that improves model answer quality, which can enhance AI utility and reliability. However, it indirectly raises concerns about over-reliance on automated self-improvement and evaluation, possibly embedding biases or errors without human oversight. Evidence is solid but effects are incremental rather than transformative.
This is an incremental technical improvement to LLM self-evaluation and refinement methodology, with measurable benchmark gains. It has modest positive implications for AI reliability and quality control but no broad societal claim beyond model performance metrics.
Evidence extracted
- CriticGen improves answer quality by 73.17% with a 93.28% non-degradation rate
- CriticGen induces higher-quality rubrics, improving relevance/coverage
SOURCE arXiv cs.AI (tier 1)
↓
DOCUMENT 165589c9-94e4-419e-b378-182d1507ae01
https://arxiv.org/abs/2609.05439
↓
EVIDENCE 2 extracted excerpts
↓
MODEL RUN 3 runs, methodology 1.0.1
↓
SCORE +34 (Favourable)
↓
CONFIDENCE 78%