Quasa
Use QUASA App
Join the pioneer of Web3 crypto freelancing today!
Open
Creator Economy

Tongue AI’s 98.71% Score Was Color Accuracy, Not Diagnosis

|Updated: |Author: QUASA Editorial Team|5 min read| 1890
Tongue AI’s 98.71% Score Was Color Accuracy, Not Diagnosis

Tongue-image AI has advanced since 2024, but the best-known 98.71% score measured color classification—not reliable diagnosis of an arbitrary illness from one photograph. Newer work targets specific conditions with larger clinical cohorts, a narrower and more defensible use of the technology.

The field is moving toward standardized imaging and decision-support systems, but the central limitation remains: visible tongue features can correlate with health conditions without uniquely identifying their cause. Current evidence supports research into screening and clinical assistance, not near-perfect home diagnosis.

What the 98.71% score actually measured

The 2024 Technologies paper describes two datasets: 5,260 images divided among seven color classes and evaluated with an 80/20 training-test split, plus 60 abnormal tongue images collected at two Iraqi hospitals from January 2022 through December 2023; XGBoost reached 98.71% accuracy on the color-classification task, compared with 91.43% for naïve Bayes, while the MATLAB interface associated detected colors with possible conditions without providing a large, independent diagnostic-accuracy trial for those diseases.

That distinction changes the meaning of the headline number. A model that assigns predefined color labels accurately has demonstrated a computer-vision capability, but it has not necessarily learned to distinguish among diseases that can share similar visible signs.

The interface used color as a bridge between an image and a list of possible conditions. It did not establish that a tongue photograph alone can determine why a color or coating is present, nor that the system can diagnose an unknown condition outside the categories represented during development.

Newer research asks narrower clinical questions

Subsequent studies have begun replacing the broad “detect illness” premise with condition-specific classification. This design lets researchers compare people with and without a defined condition and evaluate the model against a clinically identified target.

A prospective multicentre colorectal-cancer study published in 2025 used 1,389 tongue images from patients with colorectal cancer and 1,543 from participants without it; the model achieved 87.93% accuracy in internal validation, while an independent cohort of 119 cancer patients and 221 non-cancer participants produced 85.18% precision, 85% recall and an F1 score of 0.8507.

This is more clinically relevant than recognizing color because the target condition and external evaluation group are explicit. Even so, the model was framed as a complementary screening method, not a substitute for colonoscopy or a system capable of identifying unrelated illnesses.

The comparison also shows why percentages from different studies cannot be placed on one leaderboard. Color classification and colorectal-cancer screening involve different labels, populations, reference standards and consequences for false results; the lower number can therefore represent the stronger clinical experiment.

A tongue photograph is not a stable laboratory measurement

Image capture remains a major source of uncertainty. Illumination, exposure, white balance, camera hardware, viewing angle and distance can change recorded color, while food, hydration, medication and oral conditions can alter the tongue itself.

A 2025 review of computerized tongue-image analysis identifies acquisition, preprocessing, dataset construction, feature extraction and disease detection as separate sources of variation, and notes that the field still lacks unified protocols and a universally accepted large-scale standardized database.

Those gaps make results harder to reproduce and compare across hospitals, devices and patient groups. A model can perform well on images captured under one protocol yet lose accuracy when camera processing, lighting or population characteristics change.

Biology adds another ambiguity: one disease may appear through several tongue states, while one visible pattern may occur in different conditions. This is why research is increasingly combining tongue images with medical history, symptoms, saliva analysis, endoscopy, serology or established medical imaging rather than treating color as a standalone answer.

Research performance is not product authorization

A published algorithm, prototype interface or consumer app does not automatically have regulatory authorization for a medical claim. The FDA’s current AI-enabled medical-device page states that devices on its list have met applicable premarket requirements for their intended uses, while cautioning that the list itself is not comprehensive.

Authorization is tied to a particular device, manufacturer and intended indication—not to tongue-image analysis as a general technique. Evidence for one condition also cannot validate claims about diabetes, infection, anemia, cancer or other diseases collectively.

What tongue-image AI can support now

Computer vision can quantify visible features such as color, coating, texture, fissures and shape more consistently than an unaided descriptive assessment. In controlled studies, those measurements may contribute to screening, monitoring or clinical triage when they are evaluated against a defined condition and combined with other patient information.

What the evidence does not support is a general-purpose system that can diagnose whatever is wrong from a single tongue photo with near-perfect accuracy. The meaningful advance since the 2024 experiment is not universal diagnosis: it is the shift from small color-based demonstrations toward narrower clinical models, external cohorts and more explicit recognition of imaging and validation limits.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0