The Unreasonable Redundancy of Natures Protein Folds

General News

Summary

The article argues that protein sequence data is far more redundant in fold space than it appears, especially after structure prediction turns sequence scale into structural scale. It describes how the team at Ligo uses graph-based fragmentation and clustering to strip noise from predicted protein structures before training generative models for enzyme design. The analysis shows that many MGnify-derived fragments collapse into a relatively small number of structural neighborhoods, which means uniform sampling would overweight repeated folds. The piece concludes that better sampling and clustering strategies matter more than simply adding more natural sequence data if the goal is to expand usable fold diversity for design.

Classifications

industries
No industries detected
applications
Web and Content Management

AskAI Classifications

Labels
Software Development Blockchain Software Web3 Technology

Linked Companies

Edenia
$1M to $5M