Spontaneous Emergence of Symmetry in a Generative Model of Protein Structure

preprint OA: closed
📄 Open PDF Full text JSON View at publisher

Abstract

Generative models are becoming powerful tools for protein design, enabling the creation of novel protein structures and sequences. Recent approaches have shown success using diffusion models and flow matching to sample realistic protein folds. Many methods explicitly incorporate structural biases or constraints—for example, enforcing symmetry during generation—to steer the design process. Here we report the spontaneous emergence of structural symmetry in a transformer-based generative model for proteins, without any symmetry-specific conditioning during training or constraints during generation. In our flow-matching model, a single attention head in an SE(3)-equivariant transformer layer was found to be primarily responsible for the model’s ability to generate symmetric arrangements of backbone residues across chains, or repeating motifs within a chain. Our results show that protein generative models can learn high-level structural patterns implicitly from training data. This opens new questions about interpretability and control in generative design: understanding how and why a single attention head can govern a complex global property like symmetry may inform future model architectures and help exploit emergent behaviors for better protein engineering.
Full text 1,361 characters · extracted from oa-doi-fallback · click to expand
Abstract Generative models are becoming powerful tools for protein design, enabling the creation of novel protein structures and sequences. Recent approaches have shown success using diffusion models and flow matching to sample realistic protein folds. Many methods explicitly incorporate structural biases or constraints—for example, enforcing symmetry during generation—to steer the design process. Here we report the spontaneous emergence of structural symmetry in a transformer-based generative model for proteins, without any symmetry-specific conditioning during training or constraints during generation. In our flow-matching model, a single attention head in an SE(3)-equivariant transformer layer was found to be primarily responsible for the model’s ability to generate symmetric arrangements of backbone residues across chains, or repeating motifs within a chain. Our results show that protein generative models can learn high-level structural patterns implicitly from training data. This opens new questions about interpretability and control in generative design: understanding how and why a single attention head can govern a complex global property like symmetry may inform future model architectures and help exploit emergent behaviors for better protein engineering. Competing Interest Statement The authors have declared no competing interest.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00