Events

Scientists Engineer Bacteria to Survive Without One of Life's 20 Amino Acids

Researchers from Columbia and Harvard have successfully redesigned portions of E. coli to function without isoleucine, one of the 20 amino acids that form the basis of the genetic code across all known life.

7 min read
Researchers try to cut the genetic code from 20 to 19 amino acids

The genetic code underpins all life on Earth. Across nearly every organism, the same triplet sequences of DNA bases direct the production of the same 20 amino acids. This remarkable consistency, with only minor variations, suggests the code originated in the last common ancestor of all living things. Yet scientists have long wondered how this system came to be, and whether earlier organisms operated with fewer amino acids.

To explore this question, researchers at Columbia University and Harvard University set out to test whether life could function with just 19 amino acids. Their target: isoleucine, one of three structurally similar amino acids. By engineering a functional ribosome—the cellular machinery that translates genetic instructions into proteins—without isoleucine, the team demonstrated that at least some parts of the genetic code could be simplified.

Why eliminate an amino acid?

Most genetic code research has pursued the opposite direction: expanding the code to include more than 20 amino acids for novel chemical properties. This project takes a different approach, grounded in the hypothesis that early life used partial genetic codes with fewer amino acids, relying on catalytic RNAs alongside proteins to manage metabolism. While scientists have extensively studied catalytic RNAs, far less is known about what chemistry becomes possible with a reduced amino acid palette. The researchers argue that advances in artificial intelligence have made protein redesign with fewer amino acids far more feasible than in previous years.

Isoleucine stood out as a logical candidate for elimination. It belongs to a trio of similar amino acids—leucine, valine, and isoleucine—all featuring branched carbon-hydrogen structures that make them hydrophobic. These amino acids typically cluster inside proteins, away from the cell's aqueous environment. Computational analysis of the E. coli genome revealed that isoleucine was the amino acid most frequently substituted by others in related proteins across different species, suggesting it might be expendable.

Testing the concept

Rewriting all 4,500 genes in E. coli at once would be impractical and almost certainly lethal. Instead, the team began with 36 essential genes, replacing every isoleucine with valine, its closest structural relative. The results were mixed: 22 genes could not tolerate the change, but 17 survived, including one where isoleucine appeared 45 times along the amino acid chain. Even when cells survived, their growth typically slowed compared to unmodified controls—a pattern that would persist throughout the project.

Engineering the ribosome

To focus their efforts, the team targeted the ribosome itself. This massive complex of proteins and RNA molecules translates messenger RNA into proteins—essentially a hardware component needed to run a living cell from its genome. The ribosome's proteins perform critical enzymatic functions and must interact precisely with each other and with RNA molecules, making it an exacting test of whether an amino acid can be engineered out.

The researchers first swapped isoleucine for valine in 50 individual genes contributing proteins to the ribosome. Eighteen changes caused no apparent problems, 19 resulted in slower growth, and 13 were lethal. For the 32 genes showing reduced fitness, the team deployed deep-learning protein-design software to generate alternative sequences without isoleucine. Testing with four different software packages produced viable alternatives for 25 of these 32 proteins.

For the remaining five proteins, the researchers forced isoleucine changes and then used software to redesign nearby amino acids in the three-dimensional protein structure, reasoning that compensatory changes might offset structural disruptions. This approach succeeded for four of the five problem proteins.

To test whether these redesigned proteins could assemble into a functional ribosome, the team focused on the small subunit, which contains 21 proteins encoded by genes clustered together on a 10,000-base stretch of DNA. This arrangement allowed them to replace all genes in that region simultaneously.

The critical bottleneck

Working from one end of the DNA segment, the researchers successfully replaced 10 genes without problems. By 17 replacements, cell growth had slowed noticeably. At 18 replacements, the cells died. Approaching from the opposite direction revealed the same breaking point: a single gene called rplW emerged as the critical obstacle. When the team left rplW unchanged and replaced the other 20 genes, cells not only survived but grew at approximately 70 percent the rate of unmodified E. coli.

Examining the software's suggestions for rplW revealed the problem: the AI had compensated for isoleucine changes by deleting nearby amino acid sequences. While this produced a functional protein in isolation, it differed enough to fail when combined with all the other ribosomal changes.

The team resorted to exhaustive testing. They generated 16 different designs by having software suggest alternative amino acids for each of the four isoleucine positions in rplW, then tested all possible combinations. One design successfully completed the isoleucine-free small subunit, with cells growing at roughly 60 percent the rate of unmodified strains. Over 400 generations, cells accumulated 20 to 30 mutations, but none restored isoleucine to any ribosomal proteins. Notably, this modified rplW protein only functions in the context of all the other ribosomal changes; placing it alone back into the genome kills the cells.

The role of artificial intelligence

The researchers emphasize that this work would likely have been impossible without AI tools. All protein design software was AI-based, and outputs were validated using AlphaFold 2, the Nobel Prize-winning AI system for predicting protein structures. The paper highlights instances where AI made suggestions that most biologists would have rejected outright, such as replacing the flexible, neutral isoleucine with either a charged amino acid or one locked into a rigid structure.

However, the work also exposes limitations of current AI models. Unlike humans, these systems cannot explain their decision-making processes. Different models sometimes produced vastly different suggestions, which the researchers interpret as exploring different regions of sequence space—though they acknowledge this remains speculative. In at least one case, the software redesigned an entire structural element (an alpha helix) containing the modified isoleucine, for reasons the researchers cannot determine.

This underscores a fundamental reality: current AI tools enable achievements that would otherwise be impossible, yet they provide limited insight into underlying biological principles. Understanding still requires human reasoning. While developers could prioritize transparency in AI systems to expose decision-making logic, the current emphasis—reasonably—remains on producing functional results.

Achievement and uncertainty

The accomplishment is remarkable. These proteins must coordinate with each other, with ribosomal and transfer RNAs, with messenger RNA, with the proteins being synthesized, and with proteins in the ribosome's large subunit. Billions of years of evolution fine-tuned these interactions. That such radical changes could be implemented in just a couple of years is extraordinary.

The cause of the reduced growth rate remains unknown. The modified ribosome might be less accurate, producing more defective proteins through increased errors in amino acid chain assembly. Alternatively, it could be catalytically slower, constraining cell growth. Further experiments and allowing the strain to evolve over time might improve growth rates.

Whether this work provides a stepping stone toward a fully isoleucine-free genome remains uncertain. Other large protein complexes exist in cells, and some may resist AI-assisted redesign. The project's future depends on whether these research groups have the time and funding to continue. The work may ultimately tell us little about life before the universal common ancestor, given how much else in the cell has changed since then.

Yet the research may prove valuable in a different way: by inspiring other scientists to design experiments that illuminate what cells operating with a limited genetic code might actually resemble.

Source: Ars Technica · Reporting supplemented by The Silicon Ledger staff.