Opinion: Claude analyzed my genome in 30 minutes. Now we need standards for the results

Opinion: Claude analyzed my genome in 30 minutes. Now we need standards for the results

Unpacking from vacation a couple months back, I decided to give artificial intelligence a task. I paused folding my laundry, sat down with Claude, and asked it to analyze the entirety of my genome. Back in 2009, this was a task that quite literally took a village of 30 people — including me — working together to interpret the whole genome sequence of a colleague, friend, and fellow scientist named Stephen Quake. This endeavor brought together experts in genetics, computer science, pharmacology, and clinical medicine. We built an engine specifically designed for genome analysis, examined millions of genetic variants, reviewed the scientific literature, and debated which findings might matter most for his individual health circumstances. The resulting study, published in The Lancet in 2010, was among the first attempts to interpret the complete genome, and the first to translate that analysis into personalized medical recommendations. (About a year later, our team used a similar framework to analyze my own genome, though it was never officially published.) This month, Claude attempted the same task — sans the 30 people and nearly yearlong expenditure. In a prompt of 180 words, I asked Claude to follow the same framework our team had set forth back in 2009, probing it to identify rare variants that could cause an inherited disorder, genetic abnormalities that could impact my response to different medications, and whether it had enough data to estimate my risk for common diseases. Though it was a fairly simple prompt, the task was anything but. My genome file is more than a decade old, was generated using an earlier version of the human reference genome, and lacks information available in newer sequencing formats. These limitations could easily produce an analysis that is misleading or inaccurate. Despite that, Claude was able to re-identify many of the most important findings from our original analysis, including the most consequential result in my genome: I carry two copies of the APOE ε4 variant, which is associated with an increased risk of Alzheimer’s disease. It found no major gene-disrupting variant linked to a rare inherited disorder, matching the outcome reached by our original team. It also flagged variants in the DPYD and CYP2C19 genes, which can influence how we respond to certain medications. I asked it to calculate genetic risk scores for common diseases, but it rightfully warned me that my file did not contain enough information to do so reliably. Amazingly, this entire interaction used only around 400,000 tokens and cost approximately $5. Most of the half-hour was spent just waiting for scientific databases to load. All in all, a task that once took nearly a year and required a large, specialized team can now be attempted at home for roughly the price of a two-liter bottle of soda. That accessibility opens doors once thought unimaginable. And as these doors open to the general public, the scientific community has a responsibility to develop clear standards for what makes up a medically reliable genome in the first place. A genome is not a definitive record It is tempting to think of a genome sequence as a permanent record that can be preserved forever after one read. In reality though, every genome analysis is influenced by multiple of factors, such as the reference genome used for comparison and the technology used to sequence the DNA itself. Sequencing technologies have seen enormous improvement over the past couple of decades. Scientists can now examine sections of DNA that were once inaccessible and detect forms of genetic variation that older methods may have missed. But limitations remain. Almost all existing genomes were produced using a technique called “short-read sequencing,” which breaks DNA into millions of fragments and uses software to reconstruct where these fragments belong. Though this approach works well across much of the genome, it can be less reliable in the regions where sequences can repeat, resemble one another, or contain large structural changes. Though difficult to sequence, those regions are by no means irrelevant: Disease-causing variants can occur where conventional sequencing struggles to pick it up. In turn, an interpretation may report that no concerning variant was found without making clear that the relevant gene was not fully examined. Another barrier is the reference genome itself. To date, most analyses have compared patients’ genomes against a standard DNA sequence constructed from a limited number of people. As it stands, that baseline does not adequately represent the full diversity or complexity of human genomes, making some variants more difficult to detect, especially in populations that remain under-represented in genomic datasets. Newer references now represent many different versions of the human genome through branching paths rather than forcing everyone’s DNA onto one template. This can improve variant detection, especially in complex regions and across diverse populations. Yet laboratories, databases, and clinical systems have not adopted them consistently. The call for clear standards Now, with the general public able to readily use AI for genomic interpretation that may impact the course of their life and future health decisions, we have never had a more urgent need to develop standards to determine what a high-quality medical genomic interpretation should actually include. Current quality checks are efficient at detecting simple genetic changes, such as an altered DNA letter or a small piece of DNA that has been removed or added. Medically important changes, however, can be much trickier to detect. I liken it to proofreading a book: Finding a misspelled word is quite straightforward, while catching a sentence or paragraph that may not be in the right place takes a bit more nuance. A medical-grade genome test should therefore be evaluated on whether it can reliably detect the full range of genetic changes that may affect a person’s health, not only the simplest ones. This is an ask that will require a diverse set of exceptionally well-characterized genomes that laboratories, researchers, and AI developers can use as reference materials. These genomes should reflect the full sequence of both chromosome copies and the diversity of the global population, rather than relying on one person or one ancestral group as a universal standard. It is also important that the scientific community develops a shared, regularly updated catalog of medically significant genes that are historically difficult to sequence. For each gene, the catalogue would ideally include its technical complexity, which methods can reliably analyze it, and document when a specialized test is required. Standards must also reflect the task at hand. A test that is not sensitive enough can miss a finding entirely. For example, an inherited variant is present in nearly every cell, making detection fairly clear. Circulating tumor DNA, by contrast, may make up only a tiny fraction of the cell-free DNA in a blood sample. Each application therefore demands that clear minimum thresholds be set for what can be detected and how reliably, and what conclusions the result can support. The standards cannot lag behind the technology We are heading toward a world in which companies routinely give consumers access to their genomic data, including variant call files (digital files that list the genetic differences identified in a person’s DNA), that they can upload directly into systems such as Claude or ChatGPT. Consumer genetic testing was an important first step toward this kind of access, but the shift now is different in scale and interactivity. Services like 23andMe have largely focused on a relatively limited set of common genetic variants. A whole genome contains vastly more information. And with AI agents, consumers are no longer limited to a fixed interpretation of their results. They can ask follow-up questions, request additional analyses, and continue investigating particular findings. That kind of ongoing interaction with one’s own genomic data was simply not possible before. Now, imagine someone uploading their genome file into a system such as Claude or ChatGPT. Minutes later, alone at their laptop, they are told by AI that they may be at risk for Alzheimer’s disease or cancer. This may be true, but there is also the chance that the result is based on incomplete data, and the finding is much less definitive than the AI may suggest. Either way, they may feel completely blindsided, and there may not be a clinician available to talk through the result, whether it should be confirmed, or what the next steps should be. This is one of the most consequential challenges of accessible genomic interpretations. Interpretations are available before (or completely without) additional care and support. In my own experiment, Claude was able to recognize that my genome file did not contain enough information to make reliable conclusions regarding genetic risk scores for common diseases. This is encouraging — trustworthy system should be able to pump the breaks and realize when data inputs are too unreliable to make accurate estimations. Even so, that judgment should also be governed by a shared medical standard, rather than left to the discretion of AI developers. Putting that standard into practice will require coordination across specialties. Professional societies and standards organizations will need to develop and regularly update the technical requirements. Labs and AI developers should have transparent disclosures on what their systems detect reliably, and which findings still require confirmation. Health systems, meanwhile, will need practical ways to support the increasing number of patients who may arrive with AI-generated results, including confirmatory testing to genetic counseling. In 2009, the scarce resources were expertise, time, and money. Today, however, what’s scarce may be reliability. As we embrace the new bioinformatics revolution, marked by incredible speed and accessibility in data interpretation, we must urgently work towards developing standards that increase the likelihood that a casual user can upload their genomic data into an AI system and receive a trustworthy response. Euan Ashley is chair of the Department of Medicine at Stanford University and a physician-scientist whose research spans genomics, precision medicine, digital health, and artificial intelligence.

Original Source

Read the full article at Statnews →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.