Eight Letter Synthetic Genetic Code
Nature’s copy machine meets new letters
Life as we know it writes with four letters. A, T, C, and G pair up in DNA, and a tireless enzyme called RNA polymerase reads that code and copies it into RNA—the first move in turning genetic instructions into living chemistry.
What if the alphabet grew?
A team at the University of California San Diego has shown that RNA polymerase from the familiar bacterium Escherichia coli (E. coli) can accurately read and transcribe an expanded, eight-letter genetic alphabet—not only the four natural bases, but synthetic ones too. The work, published in Nature Communications, offers clear structural evidence that cells’ own molecular machinery can process information written in a richer code.
That matters because an expanded genetic language is a long-standing dream in synthetic biology: more letters mean more ways to store information, design molecules, and build systems that nature never tried.
Why four letters were never the whole story
The natural four-letter code has carried life for billions of years. It is elegant, robust, and almost universal. Still, chemists and biologists have spent years inventing extra “letters”—unnatural base pairs that fit into DNA’s double helix without collapsing it.
One well-known expanded system is sometimes called the hachimoji alphabet. It adds synthetic partners alongside A, T, C, and G. Earlier work had already shown that synthetic DNA built this way can do clever things—including molecules designed to recognize liver cancer cells. The open question was more basic and more intimate: can the cell’s own transcription enzyme treat those extra letters with the same care it gives the originals?
If the answer were no, expanded codes would stay mostly trapped in test tubes. If yes, natural machinery might one day carry synthetic information forward.
Watching the enzyme at work
Led by Dong Wang at the UC San Diego Skaggs School of Pharmacy and Pharmaceutical Sciences, the researchers put purified E. coli RNA polymerase through careful biochemical tests. They used DNA/RNA mini-scaffolds—tiny staged setups that place a single unnatural base at the critical “next letter” position on the template—and measured which nucleotide building blocks the enzyme actually added.
The unnatural bases in play included dP, dZ, dB, and dS. The team tracked incorporation of both natural and synthetic nucleotide triphosphates with gel electrophoresis and single-turnover kinetics (experiments that catch the enzyme’s first decisive step before later rounds blur the picture).
One practical snag appeared along the way: the natural base G sometimes mispaired with Z. Rather than shrug, the group synthesized a higher-fidelity cousin, Z*, swapping a nitro group for a carboxamide at a key spot on the ring. Comparing Z and Z* helped them tighten selectivity—an example of patient chemical tuning rather than a dead end.
High-resolution snapshots of recognition
Biochemistry alone could show that the enzyme worked. Structure could show how.
The team reconstituted elongation complexes—RNA polymerase caught mid-copy with either a dZ:PTP or dP:Z*TP pair in place—and solved four structures by cryo-electron microscopy (cryo-EM), a method that flash-freezes molecules and images them with electrons. Resolutions landed between about 2.42 and 2.75 Ångströms: fine enough to see the geometry of pairing and the enzyme’s grip in molecular detail.
Those maps told a reassuring story. RNA polymerase recognized the synthetic letters through the same kinds of biochemical and structural cues it uses for natural base pairs. In other words, the expanded alphabet did not force the enzyme into a totally foreign pose. It fit into the familiar choreography of transcription.
In a related study in PNAS, the same research group also reported that RNA polymerase can recognize another synthetic base pair held together without classic hydrogen bonds—the chemical “handshakes” that usually zip DNA letters shut. Hydrophobic packing and shape still helped the enzyme’s catalytic “trigger loop” close and do its job. Together, the two papers sketch a broader picture: fidelity does not always demand a perfect copy of nature’s bonding style, only a geometry the polymerase can trust.
What this unlocks—and what comes next
The immediate finding is foundational rather than clinical. These were controlled in-vitro experiments with bacterial polymerase and designed scaffolds, not full living cells rewriting their genomes overnight. That honesty is part of the hope: clear molecular rules are exactly what engineers need before bigger systems can be built safely and well.
Still, the implications fan outward. If natural RNA polymerase can faithfully transcribe eight-letter information, designers gain a sturdier bridge between synthetic DNA and the rest of biology’s toolkit. Future diagnostics, therapeutics, and engineered microbes or molecular machines could, in principle, carry extra information without inventing an entirely new transcription engine from scratch.
The liver-cancer–binding synthetic DNAs of earlier studies hint at one direction—highly specific molecular recognition. Custom compounds, novel polymers, and information-dense genetic circuits are others. Each will need more work: cellular compatibility, long-term stability, safety, and scale. The UC San Diego team’s structures and kinetics supply a molecular foundation those next steps can stand on.
A wider language, same careful reader
There is something quietly moving about the result. Evolution settled on four letters, yet the enzyme that reads them turns out to be less rigid than the alphabet it inherited. Give it well-designed partners, and it still listens.
That does not mean biology’s book has been rewritten. It means the reader is more versatile than we knew—and that versatility is an invitation. With high-resolution views of how synthetic pairs sit in the active site, chemists can refine the next generation of letters, and synthetic biologists can ask bolder questions about what living systems might one day say.
Four letters built the world we know. Eight, carefully transcribed, might help us draft a few new sentences—one faithful copy at a time.
