Google’s CoDaS Finds 66 Biomarker Candidates—but None Is Clinically Validated

Google Research’s August 21 introduction of its Biomarker Discovery Framework detailed 41 mental-health and 25 metabolic biomarker candidates generated across three cohorts totaling 9,279 participant-observations. The CoDaS-based system is designed to prioritize research hypotheses from wearable and clinical data, not diagnose patients.
The central limitation is clinical validity: none of the 66 candidates has been prospectively validated for medical use. AI Understanding’s independent assessment likewise distinguishes the candidate associations from causal evidence, regulatory clearance, external replication and demonstrated clinical utility.
CoDaS separates agent reasoning from numerical analysis

CoDaS divides biomarker research into six phases covering data understanding, literature-grounded hypothesis formation, statistical and machine-learning analysis, adversarial testing, assessment of mechanisms and novelty, and report assembly. An Orchestrator converts a researcher’s request into a plan, while specialized agents inspect datasets, propose features, challenge assumptions and examine possible explanations.
The language-model agents do not replace the core calculations. Deterministic code constructs features, estimates associations, adjusts for multiple testing and evaluates predictive performance. Shared memory and a structured fact sheet preserve the numerical record that the report-writing agents must use.
Critic and Defender agents examine leakage, overfitting, confounding sensitivity, overlap between a candidate and its target outcome, instability and physiological implausibility. AI Evolution’s account of the architecture also describes the separation of deterministic analysis, generative reasoning, an 11-test filtering stage and human oversight.
Human review remains a constraint on interpretation and follow-up. That division matters because a biologically plausible explanation generated by an agent cannot substitute for a computed association, and a statistically credible association cannot by itself establish clinical value.
Five evidence levels define what the 66 candidates mean

The CoDaS technical paper places the results in a hierarchy: an 11-check internal battery, related circadian signals across two depression cohorts, association coefficients of 0.252 and 0.126 for the highlighted sleep features, a −0.374 coefficient for a derived cardiovascular-fitness index, and cross-validated R² gains of 0.040 for depression and 0.021 for insulin resistance. It characterizes the candidates as hypothesis-generating signals rather than clinically actionable biomarkers.
- Candidate generation: The system produced and prioritized associations for mental-health and metabolic endpoints. This establishes an output from the workflow, not the reliability of every candidate.
- Internal statistical checks: Candidates were assessed for replication within the available data, stability, robustness, leakage and discriminative performance. These checks can filter weak signals but remain part of the developers’ retrospective evaluation.
- Cross-cohort convergence: The two depression datasets produced different sleep-variability features connected to circadian instability. Because their measurements, feature definitions and questionnaires differed, this is convergence around a construct rather than replication of an identical biomarker.
- Prediction gains: Selected features added modest predictive information beyond demographic variables. The gains do not establish useful clinical thresholds, calibration in practice or better patient outcomes.
- Prospective clinical validation: This stage has not occurred. The candidates have not been shown prospectively to improve screening, diagnosis, monitoring or treatment decisions.
The framework’s internal labels therefore use validation in a narrower statistical sense. A candidate may survive the battery and still fail when evaluated with another device, population, endpoint or clinical workflow.
The cohort findings remain associative

In the Digital Wellbeing cohort, variability in sleep duration was associated with PHQ-8 depression severity. In GLOBEM, sleep-onset variability had a weaker association with PHQ-4 scores and low discriminative performance. The common circadian theme is potentially useful for developing hypotheses, but it is not direct feature-level replication.
For the metabolic analysis, CoDaS constructed a cardiovascular-fitness index by dividing steps by resting heart rate and associated it with insulin resistance. The system also recovered an established relationship involving the AST-to-ALT liver-enzyme ratio, illustrating that its output can include both newly constructed wearable measures and associations already recognized in clinical research.
These coefficients describe relationships inside the analyzed cohorts. They do not demonstrate that changing sleep timing, step count or resting heart rate would alter depression severity or insulin resistance. The evaluation also does not establish diagnostic sensitivity, specificity, treatment benefit or generalization across new wearable devices and populations.
Clinical utility is the unresolved test
A prospective evaluation would need a predefined clinical purpose, such as screening, monitoring or treatment guidance. Researchers would then have to reproduce the association in independently collected data, verify measurement reliability across devices and populations, compare performance with established practice and determine whether using the candidate changes decisions or outcomes.
CoDaS may narrow a large search space and make the path from an agent-generated idea to a statistical result more traceable. Its deterministic computations, adversarial checks and human review address familiar weaknesses in automated analysis, although the current evidence does not show how consistently those safeguards will work beyond the evaluated cohorts.
As of August 24, the 66 candidates remain prioritized research hypotheses. External replication and prospective clinical evidence—not internal labels, association strength or favorable assessments of generated reports—will determine whether any becomes a clinically useful biomarker.
Subscribe to our newsletter
Get the latest Web3, AI, and crypto news delivered straight to your inbox.