Showing posts with label clinical decision instruments. Show all posts
Showing posts with label clinical decision instruments. Show all posts

Friday, October 4, 2013

Tools in the ED - Clinical Decision Instrument Basics

The Gist:  Clinical decision instruments (CDIs) are all the rage in Emergency Medicine, especially for trainees still developing gestalt; however, these tools often require proper understanding and finesse for correct utilization.  FOAM (Free Open Access Medical education) sources such as Dr. Radecki's posts on NEXUSPECARN abdominal trauma, and the Ottawa SAH Rule as well as Dr. Spiegel's post on the Ottawa SAH Rule have helped hone the way I think about and utilize decision aids. This editorial in Annals of Emergency Medicine (podcast here) is a concise, excellent synopsis of questions to ask when evaluating CDIs.

CDIs are tools, not rules.  These are typically derived through statistical methods in discrete populations. While the tools then undergo validation, these aids are artificial creations to assist providers in decision making and are not infallible. In the words of Mel Herbert regarding CDIs in Oct 2013's EMRAP : "you don't need to slavishly follow them."

What does the decision tool add to the clinical context?
  • Is the CDI better than clinician gestalt in pursuing work-up or treatment of a disease process?
    • In the editorial, Green explains this well using the PECARN blunt abdominal tool.  Physician gestalt in ordering CTs for clinically significant abdominal injury: Sensitivity 99%, Specificity 56%.  The PECARN tool offered a sensitivity of 97% and specificity of 42% [1].  Thus, no real added benefit from the tool.
    • Numerous studies investigating pulmonary embolism (PE) have determined that strict application of tools perform no better than physician gestalt within the study populations [5]. 
  • Is the tool usable? Washington University's EM Journal club covered an example of issues with usability ACS CDIs.
Clinical decision aids shouldn't replace gestalt.  
  • CDIs often appear to distill and codify components that comprise gestalt, which may be an enticing way to substitute clinical judgment.  As a medical student, I used these tools to aid in developing gestalt.  However, this could potentially be a bad habit in the making (see next point).
  • Many CDIs utilize gestalt as an entry criteria or as part of the actual aid.  
    • For example, in Tintinalli, Dr. Jeff Kline recommends applying PERC when the gestalt is there's a <15% chance that the patient has a PE, as this was the way in which the CDI was validated [3,4]. Thus, applying PERC to the wrong population may be deleterious.
  • Dr. Seth Trueger posted his PE diagnostic algorithm following an international Twitter debate on pathways and pre-test probability.  The gist of both of these is that a provider should consider the patient's clinical situation and downstream consequences or work up that may result. 

    What clinical question was the decision tool designed to answer?
    • Tools such as the Wells and Geneva scores were designed and validated as risk stratification tools, not rule-out or rule-in criteria.  
      • In this podcast, Dr. Scott Weingart offered some points of clarification on using CDIs to determine which patients to work up for PE.  He also harps on the point above - these scores are not designed to make the decision to work up/not work up a PE.
    • Measured outcome. Does the outcome reflect the clinical parameter you care about?
      • The Canadian Head CT aid seeks to identify head injuries that required neurosurgical intervention, not those that would resolve with no alteration in management.  One must decide whether this is the outcome both provider and patient care about.
    Is the patient part of the applicable population?
    • For example, it's important to note that the Canadian Head CT aid only applies to patients with: GCS 13-15, witnessed LOC, amnesia to the head injury event, or confusion and the authors excluded patients with "minor head injuries" that didn't have one of the aforementioned factors (see this post for more specific discussion of this example) [6].  Broadly applying the tool to patients who don't meet inclusion criteria or were excluded in the studied populations may lead to inappropriate stratification or intervention.
    • The performance of decision aids may depend on the prevalence of disease in the population. For example, PERC and Wells perform less well in high prevalence populations [5].
    • Various other factors such as developing vs developed setting, resources, etc may also alter the applicability of the decision aid in one's population. The more similar a paper's population is to your own, the more usable the decision aid.  For example, some decision aids may rely on a neurological exam performed by a neurologist versus an emergency physician.
    Has the decision aid been validated? If so, how?
    • Once a group derives a CDI, the tool must be validated to test it's rigor.  Dr. Newman gives a great explanation on this podcast (20 min mark). There are a few ways in which this typically happens:
      • Internal or external - validated in the same institution(s) or in other populations
      • Prospective or retrospective - data collected prospectively or retrospectively
      • Statistical or clinical - tool validated through statistical means or in "real life." The latter demonstrates usability and utility. 
      • Example: one can continue a data-collection study of parameters of the derivation portion of the study or one can use the tool in a population going forward to determine clinical utility. This is an example of the latter using the Canadian Head CT aid.
    Know whether a decision tool is a one-way or two-way instrument.  Misapplication of these tools may lead to excessive resource utilization and undermine the specificity of the aids.  [1]
    • One Way Decision Tools - Useful if all criteria are met.
      • Example: If someone is negative by PERC when utilized appropriately, it can indicate that the patient's risk of PE is below the test threshold. Conversely, one cannot say that if a patient is not PERC negative, then they necessitate work up for PE. 
    • Two Way Decision Tools - Can help a clinician decide both when to pursue an action and when not to pursue the action.  The "Ottawa ankle rule" is an example. [1]
    Note: I'm a mere novice with minimal statistics or EBM training, so these thoughts are more to be a reminder for myself than an in-depth analysis.

    References:
    1.  Green SM.  When do clinical decision rules improve patient care?  Ann Emerg Med. 2013 Aug;62(2):132-5. doi: 10.1016/j.annemergmed.2013.02.006. Epub 2013 Mar 30.
    3.  Kline, J.  Thromboembolism.  Tintinalli's Emergency Medicine.  7th ed.  p 434.
    4.  Kline, J. Prospective multicenter evaluation of the pulmonary embolism rule-out criteria. Thromb Haemost. 2008 May;6(5):772-80. doi: 10.1111/j.1538-7836.2008.02944.x. Epub 2008 Mar 3.
    5. Lucassen W, Geersing GJ, Erkens PM, et al. Clinical decision rules for excluding pulmonary embolism: a meta-analysis. Ann Intern Med. 2011 Oct 4;155(7):448-60. doi: 10.7326/0003-4819-155-7-201110040-00007.
    6.  Stiell IG, Lesiuk H, Wells GA, et al.  The Canadian CT Head Rule Study for patients with minor head injury: rationale, objectives, and methodology for phase I (derivation). Ann Emerg Med. 2001 Aug;38(2):160-9.

    Saturday, September 7, 2013

    Kappa - It's Greek to Me

    The Gist:  Many junior physicians use clinical decision instruments as an objective means of risk stratification or clinical decision making; however, these have subjective components.  Kappa, a measure of interrater agreement, is a commonly expressed statistic in medical literature, particularly in clinical decision aids.  Understanding the use, strengths, and weaknesses of kappa may help with application of decision aids and appraisal of literature.

    The Case: A 13 year old boy presented to the Janus General ED after being struck in the head with a baseball bat.  He had a slight headache, no vomiting, normal mental status, and unremarkable physical exam except a hematoma over his left parietal region.
    • I presented the case as low-risk by PECARN with ~<0.05% chance of a clinically significant injury.  An attending inquired as to how I determined that the mechanism was "not severe."  Would my assessment change if Mark McGuire swung the bat that hit my patient?  Similarly, where was my threshold with the 18 month old that fell off a bed? Did the precise number of feet matter? The truth was, probably not, not because it wasn't listed in the objective criteria of the decision aid, but because after my assessment of the patient, I already estimated that the likelihood of a clinically significant injury was minimal. I wondered:  How did they come up with these variables (was there really a difference between falls from 3 ft and 4 ft)? How frequently would other people disagree with my seemingly "objective" determinations?
    I found a paper by Nigrovic et al the next day that evaluated the agreement between nurses and physicians in the application of PECARN to mild blunt head injury pediatric patients.  This study demonstrates the differential level of agreement, or reliability, between elements of the PECARN predictors - with notable differences between subjective and objective components.*  For example, everyone agreed on vomiting, but anything containing the word "severe" was a little more nebulous.
    • History of vomiting - 97% agreement between nursing and physician assessment, with an outstanding kappa of 0.89 (95% CI 0.85-0.93). 
    • Severe injury mechanism - 76% agreed, kappa 0.24 (95% CI 0.13-0.35) in the age<2 cohort and kappa = 0.37 (95% CI 0.29-0.45) in the age 2-18 group.
    Wait, what is this kappa (k) business?
    • It quantifies interrater reliability - a measure of the degree of agreement between observers that is greater than chance alone.
      • Sometimes, even in medicine, clinicians and trainees guess.  For example, when reading a radiograph and deciding on atelectasis versus infiltrate, a physician may hedge and choose one.  This may seem straightforward, but imagine a variable such as severity of headache.  Suppose one clinician has a terrific headache and rates headaches encountered that day as non- or less severe.  The cases when that clinician and another agree would therefore be based on chance.  
    • Calculation: (Observed Agreement - Agreement Expected by Chance)/(1-Agreement Expected by Chance) - Ok, so, the actual calculation is more complicated and is explained here.
    • Assesses precision/reliability
      • Using the aforementioned study, one can see that nurses and physicians reliably detected the presence of vomiting but less reliably agreed on the presence of a severe mechanism of injury or severe headache.
    What does the value mean?
    • -1.0 = perfect disagreement, +1.0 = perfect agreement


    What are the limitations of kappa?
    • The expected agreement is affected by abnormal prevalence.  In a skewed sample, the observed agreement may be markedly different than the relative agreement (1).  This is referred to as the kappa paradox, and there are various ways to compensate for this issue.
      • Rare findings - agreement between observers may not be as reliable and will be reflected by a lower kappa.  Looking at the Nigrovic et al paper, the kappa for palpable skull fracture is abysmal at 0.00, yet the proportion of physicians and nurses in agreement was 98%.  This exists as a product of the rarity of the finding, as 1/434 and 7/434 physician and nursing assessments were positive, respectively.  Similarly, signs of basilar skull fracture was fair at 0.37 with an enormous confidence interval (95% CI 0.07-0.67).  
    • Generalizability. Diversity of skill/experience may affect kappa.
      • Are the raters emergency physicians? medical students? specialized radiologists?
      • This was ostensibly what Nigrovic et al sought to determine - do clinicians at various levels of expertise agree?  The answer - it depends.  
    What now? As a junior trainee, the ways I evaluate patients and objective data is different than that of a senior clinician.   Thus, I'm armed with this knowledge to acknowledge the limitations of the clinical decision instruments I use, understand why and how the variables are not hard and fast "rules," and use both to better patient care.
    *Note: The developers of PECARN (original study) only selected criteria with a minimum kappa of 0.5 (with a lower bound of the confidence interval of 0.40).

    References
    1.  de Vet HC, Mokkink LB, Terwee CB, Hoekstra OS, Knol DL.  Clinicians are right not to like Cohen’s κ 2013;346:f2125
    2.  Nigrovic LE, Schonfeld D, Dayan PS, Fitz BM, Mitchell SR, Kuppermann N. Nurse and Physician Agreement in the Assessment of Minor Blunt Head Trauma. Pediatrics. 2013. Available at: http://www.ncbi.nlm.nih.gov/pubmed/23979081. Accessed August 29, 2013.