Improving LLMs via Validator-to-Generator Alignment. Rodriguez, J. D., Zhang, J., Erk, K., & Durrett, G. July, 2026. arXiv:2607.02668 [cs.CL]
Paper doi abstract bibtex Large language models are inconsistent: varying prompts or including unrelated information can lead to unexpected changes in model outputs. The generator-validator (G-V) gap is one manifestation of this phenomenon, where LLMs generate responses that they then deem as invalid if re-queried to validate them. In this work, we introduce a new formulation of G-V consistency that involves a principled correction for utterance frequency. Specifically, generators often assign low likelihood to valid strings simply because those strings are a priori unlikely, which makes naive notions of G-V consistency unworkable. We show that under a natural model of rational agents answering questions with multiple answers, consistency of the validator with a frequency-corrected generator score emerges naturally. Our method, \emph\\FCPAname\ (\FCPA), is a training objective implementing frequency-corrected G-V consistency for real-world LLMs. Our experimental results show that training with \FCPA\\ substantially improves both G-V consistency and generator performance over prior methods, with gains of up to $+27$pp in Pearson correlation on IFEval and HumanEval, while preserving validator quality across all evaluated tasks.
@misc{rodriguez_improving_2026,
title = {Improving {LLMs} via {Validator}-to-{Generator} {Alignment}},
url = {http://arxiv.org/abs/2607.02668},
doi = {10.48550/arXiv.2607.02668},
abstract = {Large language models are inconsistent: varying prompts or including unrelated information can lead to unexpected changes in model outputs. The generator-validator (G-V) gap is one manifestation of this phenomenon, where LLMs generate responses that they then deem as invalid if re-queried to validate them. In this work, we introduce a new formulation of G-V consistency that involves a principled correction for utterance frequency. Specifically, generators often assign low likelihood to valid strings simply because those strings are a priori unlikely, which makes naive notions of G-V consistency unworkable. We show that under a natural model of rational agents answering questions with multiple answers, consistency of the validator with a frequency-corrected generator score emerges naturally. Our method, {\textbackslash}emph\{{\textbackslash}FCPAname\} ({\textbackslash}FCPA), is a training objective implementing frequency-corrected G-V consistency for real-world LLMs. Our experimental results show that training with {\textbackslash}FCPA\{\} substantially improves both G-V consistency and generator performance over prior methods, with gains of up to \$+27\$pp in Pearson correlation on IFEval and HumanEval, while preserving validator quality across all evaluated tasks.},
language = {en},
urldate = {2026-07-15},
publisher = {arXiv},
author = {Rodriguez, Juan Diego and Zhang, Jocelyn and Erk, Katrin and Durrett, Greg},
month = jul,
year = {2026},
note = {arXiv:2607.02668 [cs.CL]},
keywords = {Computer Science - Computation and Language, WG: Explorable},
}
Downloads: 0
{"_id":"Njy5mhgz9FCCJQeR9","bibbaseid":"rodriguez-zhang-erk-durrett-improvingllmsviavalidatortogeneratoralignment-2026","author_short":["Rodriguez, J. D.","Zhang, J.","Erk, K.","Durrett, G."],"bibdata":{"bibtype":"misc","type":"misc","title":"Improving LLMs via Validator-to-Generator Alignment","url":"http://arxiv.org/abs/2607.02668","doi":"10.48550/arXiv.2607.02668","abstract":"Large language models are inconsistent: varying prompts or including unrelated information can lead to unexpected changes in model outputs. The generator-validator (G-V) gap is one manifestation of this phenomenon, where LLMs generate responses that they then deem as invalid if re-queried to validate them. In this work, we introduce a new formulation of G-V consistency that involves a principled correction for utterance frequency. Specifically, generators often assign low likelihood to valid strings simply because those strings are a priori unlikely, which makes naive notions of G-V consistency unworkable. We show that under a natural model of rational agents answering questions with multiple answers, consistency of the validator with a frequency-corrected generator score emerges naturally. Our method, \\emph\\\\FCPAname\\ (\\FCPA), is a training objective implementing frequency-corrected G-V consistency for real-world LLMs. Our experimental results show that training with \\FCPA\\\\ substantially improves both G-V consistency and generator performance over prior methods, with gains of up to $+27$pp in Pearson correlation on IFEval and HumanEval, while preserving validator quality across all evaluated tasks.","language":"en","urldate":"2026-07-15","publisher":"arXiv","author":[{"propositions":[],"lastnames":["Rodriguez"],"firstnames":["Juan","Diego"],"suffixes":[]},{"propositions":[],"lastnames":["Zhang"],"firstnames":["Jocelyn"],"suffixes":[]},{"propositions":[],"lastnames":["Erk"],"firstnames":["Katrin"],"suffixes":[]},{"propositions":[],"lastnames":["Durrett"],"firstnames":["Greg"],"suffixes":[]}],"month":"July","year":"2026","note":"arXiv:2607.02668 [cs.CL]","keywords":"Computer Science - Computation and Language, WG: Explorable","bibtex":"@misc{rodriguez_improving_2026,\n\ttitle = {Improving {LLMs} via {Validator}-to-{Generator} {Alignment}},\n\turl = {http://arxiv.org/abs/2607.02668},\n\tdoi = {10.48550/arXiv.2607.02668},\n\tabstract = {Large language models are inconsistent: varying prompts or including unrelated information can lead to unexpected changes in model outputs. The generator-validator (G-V) gap is one manifestation of this phenomenon, where LLMs generate responses that they then deem as invalid if re-queried to validate them. In this work, we introduce a new formulation of G-V consistency that involves a principled correction for utterance frequency. Specifically, generators often assign low likelihood to valid strings simply because those strings are a priori unlikely, which makes naive notions of G-V consistency unworkable. We show that under a natural model of rational agents answering questions with multiple answers, consistency of the validator with a frequency-corrected generator score emerges naturally. Our method, {\\textbackslash}emph\\{{\\textbackslash}FCPAname\\} ({\\textbackslash}FCPA), is a training objective implementing frequency-corrected G-V consistency for real-world LLMs. Our experimental results show that training with {\\textbackslash}FCPA\\{\\} substantially improves both G-V consistency and generator performance over prior methods, with gains of up to \\$+27\\$pp in Pearson correlation on IFEval and HumanEval, while preserving validator quality across all evaluated tasks.},\n\tlanguage = {en},\n\turldate = {2026-07-15},\n\tpublisher = {arXiv},\n\tauthor = {Rodriguez, Juan Diego and Zhang, Jocelyn and Erk, Katrin and Durrett, Greg},\n\tmonth = jul,\n\tyear = {2026},\n\tnote = {arXiv:2607.02668 [cs.CL]},\n\tkeywords = {Computer Science - Computation and Language, WG: Explorable},\n}\n\n\n\n","author_short":["Rodriguez, J. D.","Zhang, J.","Erk, K.","Durrett, G."],"key":"rodriguez_improving_2026","id":"rodriguez_improving_2026","bibbaseid":"rodriguez-zhang-erk-durrett-improvingllmsviavalidatortogeneratoralignment-2026","role":"author","urls":{"Paper":"http://arxiv.org/abs/2607.02668"},"keyword":["Computer Science - Computation and Language","WG: Explorable"],"metadata":{"authorlinks":{}},"downloads":0},"bibtype":"misc","biburl":"https://bibbase.org/zotero-group/pratikmhatre/5933976","dataSources":["yJr5AAtJ5Sz3Q4WT4"],"keywords":["computer science - computation and language","wg: explorable"],"search_terms":["improving","llms","via","validator","generator","alignment","rodriguez","zhang","erk","durrett"],"title":"Improving LLMs via Validator-to-Generator Alignment","year":2026}