Every known living organism on Earth has operated with four letters: A, T, C and G. A team of scientists at the University of California San Diego have enabled cellular machinery to read twice as many.
In a system outside a living cell, the researchers showed the machinery could read the expanded alphabet, but challenges to using it inside organisms remain.
The four natural DNA letters pair up as A-T and C-G. In 2019, researchers developed four additional DNA letters: P, Z, B and S, which form two new pairs, P-Z and B-S.
Together, these eight letters create four possible base pairs, giving scientists more combinations for encoding genetic information and, potentially, more instructions.
Led by Dong Wang, the team at the university’s Skaggs School of Pharmacy and Pharmaceutical Sciences tested whether purified E. coli RNA polymerase could read the synthetic DNA base pairs during transcription.
Using cryo-electron microscopy, which captures frozen molecules at near-atomic scale, they found the enzyme transcribing DNA containing both synthetic base pairs.
The enzyme recognized the synthetic pairs using the same structural signals it uses on natural ones.
Transcription matters because DNA must be copied into RNA before a cell can use its genetic instructions.
If cellular machinery like RNA polymerase cannot recognize the expanded alphabet, the additional letters are useless.
But the enzyme’s ability to recognize DNA proved more flexible than expected.
A companion study in Proceedings of the National Academy of Sciences found that RNA polymerase could also recognize a separate synthetic base pair that holds together without the hydrogen bonds natural DNA relies on.
In both studies, researchers tested this using purified enzymes and reconstituted molecular systems rather than living cells. Integrating the expanded alphabet into cells and using it for protein translation, the step where cells build functional proteins from RNA, remains a “future goal.”
Researchers had taken a step toward expanding the genetic system in 2017, engineering a semisynthetic E. coli that carried a different unnatural pair and passed it to its daughter cells.
That strain depended on an algal transporter to import the synthetic nucleotides it could not produce on its own, leaving the system reliant on an outside supply.
That put one unnatural pair in a living cell, compared with two pairs in a test tube in 2026.
An expanded genetic alphabet could give scientists more codons, the three-letter sequences of genetic code that tell cells which building blocks to add when making proteins.
More codons could mean proteins built with amino acids that natural biology does not use. The technology could also aid diagnostics and therapeutics.
Researchers have already built strands from expanded alphabets that recognize liver cancer cells.
The outside-supply dependency that limited the 2017 strain could double as a containment mechanism, but whether this process can work outside a lab remains an open question.
