Variable-Length Generative Protein Design via Generalized Poisson Flow
2026-07-10 • Machine Learning
Machine Learning
AI summaryⓘ
The authors developed GPFlow, a new method that can generate proteins of different lengths, unlike previous models that needed a fixed length upfront. This flexibility helps explore more design options for proteins, which is important because the best protein length is often unknown. They tested GPFlow in several protein design tasks and found it worked better or as well as existing methods, especially in matching length distributions and handling complex design challenges. Their approach also comes with theoretical guarantees about its accuracy in mimicking real protein data.
protein designgenerative modelsdiffusion modelsPoisson processnegative log-likelihoodKL divergencemotif scaffoldingsequence designpeptide co-designRiemannian modalities
Authors
Chaoran Cheng, Zhanghan Ni, Yanru Qu, Yuxin Chen, Ruihan Guo, Jiajun Fan, Ge Liu
Abstract
The ability to generate variable-length proteins is crucial in protein design, where the optimal length is often unknown and tightly coupled to designability. Current diffusion- and flow-based generative models typically require the protein length to be specified before sampling, limiting their flexibility in exploring the feasible design space. To address this limitation, we introduce Generalized Poisson Flow (GPFlow), a variable-length generative framework that learns the rate function of an inhomogeneous generalized Poisson process by minimizing its negative log-likelihood. We establish population-level guarantees for recovering the joint multimodal distribution and derive an upper bound on the KL divergence between the data and generated distributions. We comprehensively evaluate GPFlow across structure and sequence design, motif scaffolding, and peptide co-design, spanning Euclidean, categorical, and Riemannian modalities to fully validate its variable-length generation quality. In unconditional design, GPFlow improves structural designability and achieves the best distributional fitness for sequence design compared to their corresponding fixed-length baselines, while perfectly recovering the length distribution. In conditional motif scaffolding, GPFlow ranks first on 10 of 16 structure-based design tasks with significantly more unique successes and also achieves more passed tasks in sequence-based design. In peptide co-design, GPFlow remains competitive even without access to a native-length oracle.