Health ArticleEducational review — not personal medical advice

Do Online Doctor Ratings Actually Tell You Anything? Study Says No

If you've ever chosen a doctor based on online reviews, a new study suggests you may want to look deeper.

14 min

Table of Contents

Key Points

  • Online doctor ratings did not predict quality, value, or peer assessments in a study of 78 specialists.
  • Among lowest-performing physicians, only 5% to 32% had low online ratings across platforms.
  • Ratings may reflect office experience or wait times, not clinical skill.
  • Patients should ask primary care doctors, check Choosing Wisely adherence, and consider hospital-level quality data.
  • Free-text comments may offer useful insights about communication and office logistics, but star ratings alone are not reliable.

Background: Why This Research Matters

We've all been there—feeling unwell, needing a specialist, and pulling out our phones to scan online reviews. It turns out, most of us are doing exactly that. A 2012 survey of 2,137 American adults found that 65% were aware of physician-rating websites, and 59% said these sites were either "somewhat important" or "very important" when choosing a doctor.

The numbers get even more striking. A survey of 1,000 surgical patients at the Mayo Clinic revealed that 81% would seek a consultation from a physician based on positive reviews alone, and 77% would refuse to see a physician based solely on negative reviews. This trend isn't just an American phenomenon—similar patterns have been reported across Europe.

Even health insurance companies are jumping on the bandwagon. Payers are now incorporating online consumer ratings into their physician search tools, suggesting that insurance companies believe these ratings are trustworthy measures of what the researchers call "ground-truth clinical performance"—meaning how good a doctor really is at their job.

But here's the critical question: Do those stars actually reflect how well a doctor practices medicine? The most popular rating websites offer no guidance on what criteria patients should use when evaluating a physician, which makes the results difficult to interpret. Some believe the ratings measure things like office environment or staff friendliness rather than clinical skill, but the limited data available suggested the opposite—or were simply inconclusive. This study was designed to settle the question with rigorous, multidimensional data.

Study Methods: How the Research Was Conducted

The research team set out to answer a straightforward question: do online physician ratings predict actual performance? To do this, they studied 78 physicians representing 8 medical and surgical specialties at Cedars-Sinai Medical Center in Los Angeles, an 886-bed tertiary referral hospital that serves as a hybrid academic-community medical center.

The physicians were originally selected from a pool of 123 eligible doctors. Researchers excluded those who weren't using the health system's electronic health record (EHR) in their ambulatory practice (30 doctors), those who received fewer than 3 peer-review survey responses (14 doctors), and those who had no online consumer rating data available (1 doctor), leaving the final group of 78.

The physician sample looked like this:

  • Specialties: Cardiology (25 doctors, 32%), Neurology (11, 14%), Nephrology (10, 13%), Obstetrics/Gynecology (9, 12%), Endocrinology (7, 9%), Otolaryngology (6, 8%), Gastroenterology (6, 8%), and General Surgery (4, 5%)
  • Race: Caucasian (50, 64%), Asian (19, 24%), other (9, 12%)
  • Gender: Male (66, 84%), female (12, 16%)
  • Practice type: Cedars-Sinai Medical Group (36, 46%), private practice (34, 44%), faculty group (8, 10%)
  • Years in practice: Median of 33.5 years (range 19–43 years)

Researchers collected consumer ratings from the five most popular online platforms according to Google Trends: Healthgrades, Vitals, Yelp, RateMDs, and UCompareHealth. Each website invites consumers to rate physicians using a 5-star scale. The team recorded each physician's average rating and the number of ratings behind that average. Median ratings across platforms ranged from 4.0 to 4.5 stars, with physicians typically receiving 4 to 9 reviews.

To measure actual clinical performance, the team used three distinct assessments:

  1. Specialty-specific performance scores: A composite metric built by specialty physicians through an intensive stakeholder process. Doctors rated each candidate metric using the modified RAND Appropriateness Method, and the top 5 metrics within each specialty were selected. These included measures of quality (like adherence to Choosing Wisely measures) and value of care (like use of generic medications and case mix–adjusted length of stay). Each metric was weighted between 0 and 100 based on relevance, and scores were calculated for fiscal years 2013–2014 by analysts who were blinded to the consumer rating results.
  2. Primary Care Physician (PCP) survey scores: Thirty-two primary care physicians rated each specialist using a 9-point modified RAND scale (from 1, "completely disagree," to 9, "completely agree") in response to the statement: "Dr. [Name] is a strong partner in value-based healthcare." To avoid order effects, specialist profiles were presented in random order unique to each survey respondent.
  3. Administrator survey scores: Twenty administrators and hospital leaders—including departmental leaders, directors of inpatient clinical programs, and personnel with broad-based knowledge of specialist performance—used the same 9-point scale and question to rate each specialist. Additional groups relevant to individual specialties (like dialysis staff for nephrologists and endoscopy staff for gastroenterologists) were also surveyed.

Statistical analysis included scatterplots with smoothing overlays, Spearman's rank correlation coefficients, and multivariable linear regression models that adjusted for number of reviews, physician specialty, race, gender, years since medical school graduation, and practice setting. Statistical significance was set at P < .05.

Key Findings: What the Data Showed

The results were remarkably consistent across every platform and every measure of performance. Here's what the researchers found:

Ratings vs. Actual Performance

Scatterplots comparing average consumer ratings with specialty-specific performance scores showed no relationship whatsoever. Across all platforms, there was either a negative or weak correlation between consumer ratings and performance scores (correlation range: −0.18 to 0.02).

After adjusting for other factors, the multivariable models confirmed no statistically significant or meaningful association between consumer ratings and performance scores. The b-coefficient range was −0.04 to 0.04 across all platforms—meaning that even a full 1-point increase in a doctor's star rating was associated with a change of less than one-twentieth of a point in their actual performance score.

Perhaps most concerning: among physicians in the lowest quartile of specialty-specific performance scores, only 5% to 32% had consumer ratings in the lowest quartile across the various platforms. In other words, the ratings system almost completely failed to identify the doctors who were performing most poorly.

No Help from Subdomain Analysis

The researchers wondered if maybe ratings aligned with specific aspects of care even if they didn't match overall performance. They tested this by looking at subdomains:

  • Quality-of-care metrics (adherence to Choosing Wisely measures and best practice alerts): No association with ratings.
  • Value-based care metrics (generic medication utilization rate and case mix–adjusted length of stay): No association with ratings.

So whether you're talking about following evidence-based guidelines, prescribing cost-effective medications, or getting patients home promptly, online stars told patients nothing about how well a doctor performed in these areas.

Ratings vs. Peer Assessments

What about what other doctors think? Ratings also failed to correlate with peer assessments:

  • PCP (primary care physician) scores: Correlation range of −0.04 to 0.07 across platforms, with b-coefficients ranging from −0.01 to 0.3 (not significant).
  • Administrator scores: Correlation range of −0.02 to 0.18, with b-coefficients ranging from −0.2 to 0.1 (not significant).

Doctors who were highly regarded by their primary care colleagues and hospital administrators were no more likely to have high online ratings than doctors with poor peer assessments.

Consistency Across Platforms

Interestingly, ratings were somewhat consistent across platforms. A physician's score on one platform significantly predicted their score on another in 5 of 10 pairwise comparisons after statistical adjustment. This suggests that the platforms may be jointly measuring something—but whatever that something is, it has nothing to do with clinical performance.

Summary of Key Numbers

  • Median specialty-specific performance score: 0.8 (IQR 0.5–0.9)
  • Median PCP score: 6.8 (IQR 5.5–7.8) out of 9
  • Median administrator score: 6.7 (IQR 5.3–7.5) out of 9
  • Median consumer ratings by platform (stars): Yelp 4.0, Healthgrades 4.2, Vitals 4.0, RateMDs 4.0, UCompareHealth 4.5

Clinical Implications: What This Means for Patients

This study provides some of the strongest evidence to date that online physician ratings, while popular and heavily used, are not valid measures of clinical performance. The researchers framed this in terms of psychometric validity—the same framework used to evaluate any questionnaire or measurement tool.

  • Face validity fails: Online platforms offer no guidance on what criteria patients should use to rate physicians, so it's unclear what a 5-star rating even means or whether it measures any explicit construct at all.
  • Content validity fails: No consistent categories of patient experience are measured across platforms.
  • Predictive and concurrent validity fail: The ratings cannot predict future physician performance and cannot distinguish between highly performing and poorly performing physicians—the two most basic functions of any useful assessment tool.

There's an important nuance here. When looking at hospitals rather than individual doctors, patient ratings have shown some ability to predict institutional quality. Two large retrospective studies found that web-based patient ratings of hospitals predicted hard outcomes like Hospital Consumer Assessment of Healthcare Providers and Systems (HCAHPS) survey scores, as well as hospital readmissions for myocardial infarction (heart attack), heart failure, and pneumonia. But translating that institutional-level validity down to individual physicians simply doesn't work.

One study using UK National Health Service Choices data found that among the 69% of physicians who were "recommended" in survey reviews, there were moderate associations with patient experience (Spearman's correlation 0.37–0.48, P < .001) but only weak associations with clinical process and outcome measures (Spearman's less than 0.18, P < .001). Other studies found that online reviews correlated with board certification status, quality of medical education, and physician volume—but none had robustly characterized performance across quality, value, and peer review simultaneously like this study did.

The report directly challenges the approximately 80% of healthcare consumers who currently select physicians based on ratings alone. And it pushes back on insurance companies that have begun incorporating consumer ratings into physician search tools, suggesting these organizations may be relying on unreliable data.

Study Limitations: What This Research Couldn't Prove

No study is perfect, and the authors were transparent about their study's limitations:

  • Sample size and location: The study included only 78 physicians from a single health system in Los Angeles, which may limit how broadly the findings can be generalized to other regions, hospitals, or practice settings.
  • Limited performance score variation: Because the study was conducted within one health system, there may have been a restricted range of performance scores, potentially making it harder to detect associations.
  • Academic center setting: The research was conducted at a hybrid academic-community hospital, which may not capture the full spectrum of practice environments. The authors noted this avoids some sampling bias associated with purely academic settings but still limits generalizability.
  • No access to HCAHPS or direct patient satisfaction scores: The researchers did not have data from the Consumer Assessment of Healthcare Providers and Systems (CAHPS) surveys or other formal patient satisfaction instruments. While they found modest associations between online reviews and National Committee for Quality Assurance consumer satisfaction scores for insurance plans (Pearson correlation = 0.376) in prior literature, they could not directly test what latent construct the online ratings were actually capturing.
  • Small numbers of reviews: With median review counts of only 4 to 9 per physician across platforms, individual physician ratings may be based on very limited samples of patient opinions.

The researchers also noted that what the ratings are measuring remains an open question—it may partially reflect patient satisfaction, but this study could not definitively identify the latent construct.

Future Directions: How Online Ratings Could Be Improved

Despite their negative findings, the authors emphasized that patient assessments of physicians are important and will remain a permanent part of healthcare. The question is how to make them more useful. They offered three concrete recommendations:

  1. Create a definable construct. Rating websites should clearly define what their ratings are intended to measure and provide instructions on how patients should evaluate physicians across metrics they are well positioned to assess—like communication, wait times, and whether the doctor listens.
  2. Combine ratings with quality and value data. Since online ratings clearly don't capture clinical quality or cost-efficiency, they should be paired with complementary data on these critical metrics so patients can see the full picture.
  3. Acknowledge limitations. Companies offering physician rating services should take responsibility for helping consumers understand what their scores can and cannot tell them, recognizing that the stakes are high when patients are making healthcare decisions.

Recommendations: Practical Advice for Patients

So what should you do the next time you need to find a specialist? Based on this research and the broader literature, consider these strategies:

  • Don't rely on star ratings alone. This study shows that stars don't predict clinical quality, value of care, or peer reputation. Treat them as one small piece of information, not a decisive factor.
  • Ask your primary care doctor. Your PCP works with specialists every day and knows which ones communicate well, manage complex cases effectively, and provide cost-conscious, high-quality care. This is exactly the kind of expert peer assessment that online ratings failed to capture.
  • Look for "Choosing Wisely" adherence. The Choosing Wisely campaign promotes evidence-based care by helping physicians and patients avoid unnecessary tests and procedures. Doctors who follow these recommendations are practicing high-quality, value-based medicine.
  • Check hospital quality data when available. While individual physician ratings don't predict performance, hospital-level patient ratings have shown some ability to predict institutional outcomes like readmissions for heart attack, heart failure, and pneumonia (118). National hospital comparison tools from the Centers for Medicare & Medicaid Services (CMS) can provide useful context.
  • Read the comments, not just the stars. Free-text comments may give you useful insights about communication style, office logistics, and interpersonal care—aspects of the patient experience that are genuinely valuable, even if they don't measure clinical performance.
  • Seek out board certification and volume information. These were among the few factors shown to correlate with online ratings and are generally considered markers of physician quality—though remember that even these have limitations.

The bottom line from the study's authors is simple and direct: "Online consumer ratings should not be used in isolation to select physicians, given their poor association with clinical performance." Healthcare decisions are too important to be made on the basis of a star rating that may reflect little more than whether a patient liked the office furniture or waited an extra fifteen minutes.

Frequently Asked Questions

Do online doctor ratings predict the quality of care a doctor provides?

No, according to a study of 78 specialists, online star ratings showed no meaningful relationship with objective measures of quality, value, or peer assessments. Among the lowest-performing doctors, only 5% to 32% had low online ratings, so the reviews usually failed to identify those with poor clinical performance.

Can I use online reviews to choose a specialist?

The study authors advise against using online consumer ratings alone to select a physician. Ratings did not predict clinical quality, cost-efficiency, or reputation among peers. Instead, ask your primary care doctor for a referral, look for board certification, and read the free-text comments for useful details about communication and office experience.

Why don't online doctor ratings match actual performance?

The study found that online platforms offer no guidance on what patients should rate, so the results are hard to interpret. Ratings may reflect things like office friendliness or wait times rather than clinical skill. Across several platforms, correlation with performance measures was negligible, and even peer assessments did not align with stars.

What should I do instead of relying on star ratings?

Ask your primary care doctor, who works with specialists regularly and knows their communication and quality of care. Look for doctors who follow Choosing Wisely recommendations to avoid unnecessary tests, and check hospital-level quality data, which may be more informative than individual physician ratings.

Are hospital patient ratings more reliable than doctor ratings?

The article notes that patient ratings of hospitals have shown some ability to predict institutional outcomes, such as readmissions for heart attack, heart failure, and pneumonia. However, this validity does not translate to individual physicians, where online ratings failed to predict clinical performance.

Is there anything useful in online reviews, like patient comments?

Free-text comments may give useful insights about communication style, office logistics, and interpersonal care, which are valuable aspects of patient experience. However, the star ratings themselves do not measure clinical quality, value, or peer reputation, so treat comments as one small piece of information.

Source Information

Original article title: Online physician ratings fail to predict actual performance on measures of quality, value, and peer review

Authors: Timothy J. Daskivich, Justin Houman, Garth Fuller, Jeanne T. Black, Hyung L. Kim, and Brennan Spiegel

Journal: Journal of the American Medical Informatics Association (JAMIA), Volume 25, Issue 4, 2018, pages 401–407

DOI: 10.1093/jamia/ocx083

Publication dates: Received March 13, 2017; Revised May 18, 2017; Accepted August 21, 2017; Advance Access publication September 8, 2017

Funding/Institutional review: The study was approved by the Cedars-Sinai Institutional Review Board (IRB) and was conducted as part of a larger initiative within the Cedars-Sinai health system to perform quality ratings of specialty physicians for the purpose of developing new provider narrow networks.

This patient-friendly article is based on peer-reviewed research. It is intended for educational purposes and does not constitute medical advice. Always consult with a qualified healthcare professional regarding medical decisions.