Contents
pdf Download PDF
pdf Download XML
49 Views
30 Downloads
Share this article
Research Article | Volume 18 Issue 8 (AUGUST, 2026) | Pages 64 - 75
Accuracy of Alvarado Score in Diagnosing Acute Appendicitis: A Validation Study in the Emergency Departments of Islamabad's Tertiary Care Centers
Under a Creative Commons license
Open Access
Received
July 1, 2026
Revised
July 22, 2026
Accepted
Aug. 4, 2026
Published
Aug. 6, 2026
Abstract

Background: Acute appendicitis is among the most frequent conditions causing acute abdominal pain that needs an immediate surgical intervention. Prompt and accurate diagnosis of the condition is very important for avoiding unneeded operations and complications of treatment delay. The Alvarado score is a very simple and popular clinical score system used to help diagnose acute appendicitis. At the same time, the accuracy of this score may vary in various population groups. The objective of the study was to assess the diagnostic accuracy of the Alvarado scoring system based on histological criteria. Methods A cross-sectional study design was used on 300 patients with a diagnosis of clinically suspected acute appendicitis by undergoing an appendectomy. Demographic data, clinical data, laboratory data, imaging data, intraoperative findings, and histopathologic diagnosis were obtained and analyzed through SPSS version 26 software. Comparison between continuous variables was done using independent samples t-test or Mann Whitney U-test while categorical variables were analyzed by Chi square test. Binary logistic regression was done for predictive ability of Alvarado score for histopathological diagnosis of acute appendicitis. Diagnostic accuracy was done by analyzing sensitivity, specificity, PPV, NPV, accuracy, and ROC curve analysis. Results: Of the total of 300 patients enrolled in this study, 191 patients (63.7%) had histopathologically proved acute appendicitis, while 109 patients (36.3%) were found to have no positive histopathological results. Migratory pain (P = 0.013), anorexia (P < 0.001), nausea and vomiting (P = 0.026), rebound tenderness (P < 0.001), and Alvarado risk category (P < 0.001) were significantly correlated with histopathologically proved acute appendicitis. However, there were no significant correlations for gender (P = 1.000), residence (P = 0.266), ultrasound results (P = 0.690), CT scan results (P = 0.648), type of surgery (P = 0.097), and intraoperative results (P = 0.751). Age was also not significantly different between histopathology positive and histopathology negative groups (independent-samples t-test: P = 0.967; Mann-Whitney U test: P = 0.944). The binary logistic regression revealed that the Alvarado score was significantly predicting histopathologically proved acute appendicitis (β = 1.585; OR = 4.88, 95% CI: 3.36-7.09; P < 0.001). With an Alvarado score cut-off value of ≥7, the score system attained a sensitivity of 83.77%, specificity of 100.0. Conclusion: Alvarado scoring was highly accurate in diagnosing acute appendicitis confirmed by histopathology. An Alvarado score of ≥7 was an excellent indicator of high specificity and positive predictive value in addition to sensitivity and accuracy. It was a good predictor of acute appendicitis independently and can be used in the evaluation of acute appendicitis.

Keywords
INTRODUCTION

The condition of severe appendicitis still serves as one of the main reasons for performing emergency laparotomy and laparoscopy in patients of any age, and it is one of the most prevalent conditions causing severe abdominal pain, which requires urgent surgery [1]. It still presents with many diagnostic difficulties, since its presentation is often similar to many other causes of acute abdomen, including those originating from gynecology, urology, and gastroenterology, despite the current advances in imaging and laboratory testing [1,2].

 

Statistics of global epidemiology demonstrate the persisting problem of appendicitis among major diseases. The global age-standardized prevalence of 214 per 100,000 people or about 17 million cases each year, with the highest rates of incidence recorded in high-income Asia-Pacific regions [17]. Even though mortality and disability-adjusted life-years caused by appendicitis have been declining consistently since 1990, almost half of the regions included in the study have witnessed the increase in age-standardized rates of incidence [17,18]. Approximately 4.53 million new cases in children and teenagers have been diagnosed in 2021, with the number likely to rise more than 21% by 2040 [18].

 

The possibility of delayed diagnosis of appendicitis results in increased risks of perforation, peritonitis, and morbidity related to the surgery performed, but the problem of unnecessary surgery associated with overdiagnosis of appendicitis due to purely clinical suspicions results in negative appendectomies, which may even reach up to 15-20% of all procedures done [8,14]. Thus, balancing the risks mentioned above gave rise to the development of various prediction models that would facilitate decision making and reduce the impact of the subjective assessment.

 

The Alvarado Score created by Alfredo Alvarado in 1986 is the most extensively studied and widely applied clinical scoring system in the management of acute appendicitis [1]. The ten-point system is based on three symptoms (pain migrating into right iliac fossa, anorexia, nausea/vomiting), three signs (tenderness in the right iliac fossa, rebound tenderness, fever), and two laboratory data (leukocytosis and left shift/neutrophilia) [1]. The reasons for its wide acceptance are the simplicity of use, low costs, and application at the bedside without need for any imaging studies, which is particularly useful in cases of limited resources of many emergency rooms in Pakistan.

 

Nonetheless, several studies have found that the efficacy of using the Alvarado score in diagnosing appendicitis varies across populations. For instance, a systematic review involving data from 42 studies found that although the cut-off of 5 had good sensitivity in ruling out appendicitis, the cut-off of 7 (which is normally the cut-off value used to warrant an operation), had low specificity, with pooled specificity values of 57% in men and 73% in women, with the score having a propensity to over-diagnose appendicitis among women and children [2]. The differences have led to the development of other scoring systems, including the Raja Isteri Pengiran Anak Saleha Appendicitis (RIPASA) score [3] and Appendicitis Inflammatory Response (AIR) score [4], which have shown to be more sensitive and specific than the Alvarado score in Asian populations.

 

Recent comparative studies performed in Pakistan, especially in Peshawar, Lahore, Karachi, Rawalpindi and Bahawalpur, have revealed marked variability of the Alvarado score performance with sensitivities ranging between about 54% and 94% and specificities between 57% and 88% based on patient selection criteria, cut-offs and reference standards [5,7,8,11,13,15,16]. Variability highlights the need for local validation data instead of unjustified use of results obtained in other populations due to dietary differences, different health seeking behaviors, practice of self-medication with antibiotics and delays before presentation – all of which may affect clinical and laboratory signs and symptoms during the evaluation.

 

There are several big tertiary teaching hospitals in the federal capital city of Islamabad, which serve a heterogeneous population from urban, peri-urban and referred rural population from the twin cities and north Punjab/Khyber Pakhtunkhwa area. Despite high prevalence of appendicitis among patients attending their emergency rooms, there is no data available on validation of the Alvarado score in these tertiary centers in Islamabad. The study intends to fill this gap and perform prospective validation of the diagnostic accuracy of the Alvarado score with histopathology serving as a gold standard.

 

2. Literature Review

2.1 The Alvarado score and its diagnostic performance

The cut-off described by Alvarado in 1986 was that ≥7 indicated definite appendicitis, ≤4 ruled it out, and 5-6 was indeterminate, requiring monitoring [1]. In the past years, numerous validation studies done in various patient populations have examined this cut-off. The systematic review done by Ohle et al. that encompassed 42 such studies found that the score performed extremely well as a 'rule-out' test at the lower cut-off (with sensitivity of nearly 100 percent at a score of 5), but not as well as a 'rule-in' test at the higher cut-off, with decreasing specificity, particularly in females, where gynecological presentations of appendicitis are common [2].

 

This trend is also exhibited by more recent studies involving single centers. For instance, a retrospective study conducted in Yemen revealed that the sensitivity of the Alvarado score was 94.6%, and its specificity was 87.8% at a cut-off of 6, with an area under the curve of 0.985 and compared with histology and was equivalent to abdominal ultrasonography [6]. On the other hand, a previous prospective study in Jordan indicated that there was a significantly low sensitivity of 54% and specificity of 75%, and this means that the Alvarado score was not sensitive enough in their setting [8]. An Iraqi study conducted in 2025, focusing only on the ambiguous 5-7 score range, showed that the negative predictive value was low in this mid-range (45.7%).

 

2.2 Comparative studies against RIPASA and AIR scores

Several organizations have challenged the applicability of the original Alvarado group with respect to South and Southeast Asia since the former is a part of Western community. Thus, the RIPASA score was designed in Brunei considering additional features such as patient nationality/ethnicity and duration of symptoms [3]. Since then, several comparative research have been done in Pakistan where the two scales are used simultaneously. Though the Alvarado scale was slightly more accurate than the RIPASA one (89% vs 88%), the latter showed better specificity for that group of patients, according to the 2025 study done in Hayatabad Medical Complex, Peshawar [5]. On the other hand, a comparative study made in Lahore and published in the Journal of the College of Physicians and Surgeons Pakistan and a cross-sectional study from Karachi using histopathology as the gold standard confirmed the fact that RIPASA demonstrated better sensitivity for the diagnosis of appendicitis, though at the expense of specificity [7,13]. Comparable results on setting-related differences in diagnostic efficacy were found in the study conducted in Combined Military Hospital, Rawalpindi [16].

 

In Pakistani cohorts, there is also a direct comparison of the AIR score, where the laboratory signs of inflammation (neutrophils percentage and CRP) have more importance compared to the Alvarado score. None of the scores has demonstrated consistency regarding superiority over others, and local validation remains important, as concluded in a study from Hayatabad Medical Complex, comparing the Alvarado, AIR, and RIPASA scoring systems in 132 patients [11].

 

2.3 Modified and combined scoring approaches

Considering the well-known drawbacks of Alvarado score as a single diagnostic method, studies have been done on using this tool in combination with other tests. An Indian study published in 2025 used the Modified Alvarado Score in combination with ultrasound in a tertiary hospital and reported better results compared to any test used individually [12]. Likewise, a retrospective study conducted in Turkey came up with another score based on Alvarado score by incorporating imaging and laboratory tests to distinguish complicated cases of appendicitis from uncomplicated ones [10]. Another quasi-experimental study was done at Combined Military Hospital, Bahawalpur using both Modified Alvarado and RIPASA scores in a young adult group [15].

 

2.4 Emerging machine learning and artificial intelligence approaches

The application of machine learning (ML) and artificial intelligence (AI) to the diagnosis and estimation of appendicitis severity has gained momentum and quickly gone past the conventional scores. For example, in the study including over 1,800 individuals, Akbulut et al. proposed an explainable ML algorithm via CatBoost to differentiate between perforated and non-perforated appendicitis with acceptable prediction accuracy while being clinically interpretable [20]. The comparison of six machine learning algorithms to predict appendicitis severity in 2024 revealed that such systems work more effectively than grading approaches [21]. Even though the populations, input data, and outcome measures were different making any comparisons difficult and preventing generalizability to practical ED settings, systematic reviews provided some summaries of these findings noting that several ML models outperformed the Alvarado score in benchmarking [22–25]. Meanwhile, lowering the rates of negative appendectomies among the children who are at high risk of appendicitis according to pre-test likelihood has been studied in other recent papers [26, 27]. Though new methods are promising, their use in practical ED settings with limited resources remains impossible because they require computational means and external validation; therefore, Alvarado's score is still the best option [23, 24].

 

2.5 Rationale for the present study

Taken together, the literature indicates that (a) the Alvarado score's diagnostic accuracy differs greatly based on the population studied, its sex and age group, as well as the cut-off value used; (b) other scores, like the RIPASA or the AIR score, can perform better than the Alvarado score in certain South Asian populations, though they have not proven uniformly superior to it; and (c) there are only a few validation studies performed locally in the Islamabad's tertiary care emergency settings, even though there have been many studies conducted in Pakistan [5,7,11,13,15,16]. Validation research that takes place in a local setting and generates context-relevant estimates of diagnostic accuracy that could guide clinical decision-making is needed to understand the impact of such a score on triaging procedures and resource utilization in Islamabad's emergency departments.

 

3. Objectives

3.1 Primary objective

Using the cut-off point of ≥7, calculate the diagnostic efficiency of the Alvarado score in terms of sensitivity, specificity, PPV, NPV, and overall accuracy for diagnosing acute appendicitis in patients admitted to the ERs of tertiary care hospitals in Islamabad using the histopathology of the resected appendix as the gold standard.

 

3.2 Secondary objectives

● To identify the diagnostic efficiency of the Alvarado score in the low (1–4), intermediate (5–6), and high (7–10) probability categories.

● To estimate the diagnostic efficiency of the Alvarado score in children, adults, and the elderly, as well as in males and females.

● To identify the frequency of undiagnosed or late diagnoses of appendicitis and the rate of negative appendectomy for the Alvarado score.

● To construct the ROC curve and establish the locally optimal cut-off score maximizing the sum of sensitivity and specificity of the population.

●      To identify the frequency distribution of each item of the Alvarado score among cases of histologically proved and unproved appendicitis.

 

4. Research Hypothesis

4.1 Null hypothesis (H0)

In respect of acute appendicitis diagnosis in relation to histopathology in those who present themselves to the emergency departments of tertiary care centres in Islamabad, the Alvarado score is inaccurate in diagnosis even at the traditional cut-off point of ≥7 due to its inadequate sensitivity, specificity, and accuracy (which should be ≥80% in each case).

 

4.2 Alternative hypothesis (H1)

As compared to histopathological diagnosis, the Alvarado score, when applied with a cut-off point of ≥7, is relatively accurate in the diagnosis of appendicitis in patients presenting in the emergency departments of the tertiary care institutions in Islamabad (diagnostic sensitivity and specificity each ≥80% and overall accuracy ≥80%).

 

Based on previous studies, a secondary hypothesis is that there will be marked variation in diagnostic accuracy depending on sex and age group, with poor specificity amongst females in their reproductive years [2,5,7].

MATERIALS AND METHODS
5.1 Study design Prospective, cross-sectional and analytical validation study based on STARD (Standards for Reporting of Diagnostic Accuracy Studies) Guidelines. 5.2 Study setting and duration The study will take place in the general surgery wards and emergency departments of two or three tertiary teaching hospitals in Islamabad. A period of twelve months for conducting research is recommended to allow enough time for recruitment of participants, variation due to seasons in presentation, and report of histopathology for each participant recruited. 5.3 Study population Eligible participants shall be all patients irrespective of gender visiting the emergency departments participating in the study period having clinical presentation of acute appendicitis. These patients shall be presenting with pain in right iliac fossa or migratory abdominal pain with or without associated gastrointestinal or systemic symptoms. 5.4 Sample size calculation The minimum sample size as per diagnostic accuracy sample size calculation formula is approximately 190-210 patients based on anticipated sensitivity of 85% (calculated from pooled estimates in the region [2,5,7,11]), anticipated proportion of confirmed appendicitis of 70% among clinically suspected cases, 95% confidence limit, and margin of error of 6%. A goal of at least 230 patients is recommended, considering an anticipated 10% dropout (incomplete histopathology, conservative treatment without surgery, or withdrawing consent). Consecutive sampling techniques will be employed to minimize selection bias. 5.5 Inclusion criteria ● Adult patients 12 years old and above come to the hospital ER with a clinical diagnosis of acute appendicitis. ● Individuals who are willing to give their written consent to participate (or minors who are willing to go along with parental consent). ● After the surgery, either open or laparoscopic appendectomy, the removed tissue will be subjected to histopathologic examination. 5.6 Exclusion criteria ● Individuals in whom, on initial assessment, a different clinical diagnosis was readily identifiable (verified ovarian pathology, ectopic pregnancy, perforated peptic ulcer). ● Individuals receiving non-surgical management of their periappendiceal mass or abscess (interval appendectomy cases as histology cannot be obtained immediately). ● Pregnant individuals owing to the recognized effects on laboratory and clinical parameters associated with pregnancy. ● Immunosuppressed individuals or those on long-term corticosteroid therapy may exhibit a poor leukocyte reaction or fever. ● Individuals decline an appendectomy or consent to one where there is no histological specimen available for examination. 5.7 Sampling technique All patients that qualify for participation during the research study period will undergo recruitment through screening of each eligible patient until the required sample size is attained.
Recommended Articles
Research Article
Accuracy of Alvarado Score in Diagnosing Acute Appendicitis: A Validation Study in the Emergency Departments of Islamabad's Tertiary Care Centers
...
Published: 06/08/2026
Chat on WhatsApp
© Copyright CME Journal Geriatric Medicine