Introduction

Intracranial hemorrhage (ICH), or bleeding within the brain parenchyma or associated compartments, is a major global health concern1. In 2021, there were 3.44 million new cases worldwide, with an age-standardized prevalence of 40.8 per 100,000, contributing to over 3.3 million deaths and nearly half of all stroke-related DALYs2.Prompt detection is vital, since timely intervention can mitigate hematoma expansion and improve outcomes3. Non-contrast CT (NCCT) remains the first-line imaging modality due to its speed, accessibility, and diagnostic accuracy4. The clinical workload for radiologists has grown substantially in recent years, with on-call volume at one large academic center quadrupling between 2006 and 2020, largely due to increased CT use5,6.

Deep-learning models for ICH detection on NCCT show promising performance; a meta-analysis reported pooled sensitivity of 92% and specificity of 94%7. Despite FDA clearance of at least 15 models for ICH detection8, there is little to no evidence of rigorous external validation on large, demographically and technically diverse datasets, limiting their generalizability and clinical applicability in real-world settings7. For example, a systematic review of external validation studies of radiology AI models found that 81% of them demonstrated degraded performance on independent datasets9. This can be attributed to differences in patient populations, disease prevalence or presentation, and technical CT-scan parameters10.

In this study, we assess the performance of a commercial, FDA-cleared CNN model for ICH triage (Aidoc Medical, NY). The model functions by detecting positive cases of ICH from the radiologist worklist and alerting the radiologist to its presence so it may be interpreted sooner, with reported improvements in exam turnaround times11. The FDA-clearance reports sensitivity of 96.15% (95%CI: 90.44–98.94%) and specificity of 94.83% (95%CI: 89.08–98.08%) on 220 cases, and does not provide information on size, location, or acuity of hemorrhages12.

Previous studies show substantial variability in this model’s performance, with sensitivity ranging from 68.2% to 89% and specificity from 96 to 98.1%, depending on clinical context and evaluation criteria13,14,15,16. However, a limitation of existing studies is the lack of evaluation across different care settings (inpatient, outpatient, emergency) and ICH subtypes. To realize the true potential of AI in radiology and healthcare, models should augment physician performance in subtle, easily missed, or misdiagnosed cases rather than in those that are readily detected by humans. Therefore, in this study we retrospectively evaluated the model’s performance on two years of data from a large academic center hospitals across multiple demographic, clinical, and anatomical subgroups, allowing for a detailed characterization of the model’s behavior.

Results

Patients Cohort

A total of 146,226 head NCCT examinations were initially identified. After excluding inpatient post-operative exams, 103,822 studies remained. Following further exclusion of studies with indeterminate ICH diagnosis, the final cohort included 101,944 unique CT studies from 74,142 unique patients (Fig. 1). Table 1 presents the demographic distribution of the cohort.

Fig. 1: Cohort selection for head CT examinations.
Fig. 1: Cohort selection for head CT examinations.
Full size image

Selection of eligible head CT exams interpreted by the AI model, following exclusion of inpatient post-operative exams and reports mentioning possible or indeterminate hemorrhage.

Table 1 Demographic distribution of included cohort

LLM Performance for ICH Label Extraction

LLM showed an overall accuracy of 96% and a Cohen’s kappa of 0.85. The confusion matrix demonstrated excellent agreement with the expert reference standard for both positive and negative classes, with most discrepancies occurring in the possible category (Figure S1). When excluding the ‘possible’ class, LLM achieved a sensitivity of 96.5%, specificity of 98.8%, accuracy of 98.6%, PPV of 91.7%, and NPV of 99.5% for ICH detection. Table S3 shows the distribution of “possible” labels in the whole cohort of this study compared with the one-year period preceding this cohort.

ICH Subgroup Label Extraction

Detailed performance metrics for each label are summarized in Table S4. Performance for extracting hemorrhage acuity was good with sensitivity ranging from 82.9 to 100% and specificity ranging from 95.9 to 99.3%. Performance for extracting anatomical compartment was excellent with sensitivity of 98.4 to 100% and specificity of 95.2 to 99.1%. For the detection of mass effect, the LLM achieved a sensitivity of 99.1% and a specificity of 92.8%.

Performance for extracting hemorrhage location was varied but generally good. Extraction of parenchymal location demonstrated sensitivity of 83.3 to 100% and specificity of 95.4 to 100%. For extracting locations of epidural/subdural hemorrhages (EDH/SDH), sensitivity ranged from 98.5 to 98.7% and specificity ranged from 74.2% to 100%. The sensitivity ranged from 95.7% to 100% for subarachnoid hemorrhage, while specificity ranged from 71.1 to 99.1%. The LLM achieved high performance across all ventricular compartments—sensitivity ranged from 96.4 to 100%, while specificity ranged from 97.8 to 100%.

The LLM demonstrated robust performance in extracting hemorrhage size categorized into three groups: ≤10 mm, >10 mm, and “not-mentioned”. For EDH/SDH, the model achieved an accuracy of 94.7% and a Cohen’s kappa of 0.925. For intraparenchymal hemorrhage (IPH), accuracy was 95.9% with a Cohen’s kappa of 0.924.

Based on the extracted reference standard labels, 6814 examinations (6.7%) were positive for ICH (3009 patients). Acute ICH accounted for 53.6% of cases, while chronic and subacute forms were less common (3.8% and 2.9%, respectively). Acuity was not documented in 39.7% of reports, frequently using terms, such as ‘unchanged’, ‘stable’, or ‘redemonstrated’. Hemorrhage was localized to a single-compartment in 64.8% of cases and involved multi-compartments in 35.2%. Size was reported as ≤10 mm in 26.3% of cases and >10 mm in 31.5, with 42.2% lacking size documentation (usually mentioned as ‘small’, ‘tiny’, or related to subarachnoid and intraventricular). Mass effect (including midline-shift) was observed in 55.6% of examinations. Prevalence of all labels varied significantly across patient class, with additional differences observed across race, age, and sex. Full distributions are provided in Tables S5 and S6.

Model Performance for ICH Detection

The model demonstrated an overall sensitivity of 82.2% and specificity of 97.6% for ICH detection across all exams with an overall accuracy of 96.6%, PPV of 70.8%, NPV of 98.7%, and an F1 score of 76.1%. Table S7 shows sensitivity analysis of the model performance treating “possible” cases as ICH positive or negative.

Demographic subgroup performance

Table 2 provides full demographic subgroup performance analysis. Model sensitivity was lowest in outpatients (72.2%, CI: 68.4.4–76.1) compared with inpatients (82.7%, CI: 81.7–83.8) and ED patients (83.5%, CI: 81.8–85.1). Sensitivity was slightly lower among Black patients (80.7%, CI: 79.4–82.0). By sex, sensitivity was 82.8% (CI: 81.5–84.1) in females and 81.7% (CI: 80.4–82.9) in males. It was lower in patients under 40 (79.2%, CI: 75.9–82.4) compared with those over 65 (82.9%, CI: 81.7–84.1). Specificity was consistently high across subgroups, but was notably higher in ED patients (98.3%, CI: 98.2–98.4), females (97.8%, CI: 97.7–98.0), and those younger than 40 (98.8%, CI: 98.6–98.9).

Table 2 Aidoc model performance across demographic subgroups

Hemorrhage Subtype Performance

Table. 3 summarizes model sensitivity across ICH subgroups, highlighting variability by acuity, hemorrhage size, anatomical compartment, and location.

Table 3 Model performance across different ICH subgroups

The model demonstrated highest sensitivity for acute (86.2%, CI: 85.1–87.3) and acute–on–chronic (89.6%, CI: 85.7–93.3) hemorrhages, with large drop-offs in performance for chronic (54.8%, CI: 49.0–60.9) and subacute (45.5%, CI: 38.3–52.0) hemorrhages. Further, sensitivity was significantly lower for single-compartment hemorrhages (76.0%, CI: 74.7–77.2) compared to two-compartment (91.9%, CI: 90.5–93.2) and >2 compartments (96.9%, CI: 95.5–98.0) hemorrhages.

For single-compartment cases, sensitivity was consistent across locations, 76.0% (CI: 74.6–77.5) for all extra-axial, 76.3% (CI: 74.5–78.1) for EDH/SDH, and 74.4% (CI: 70.6–78.4) for SAH. Within both the subdural/epidural and subarachnoid compartments, diffuse hemorrhages (involving multiple locations inside the compartment) showed the highest sensitivity: 89.4% (CI: 86.6–92.2) and 88.4% (CI: 80.8–94.9), respectively (Table S8).

Intra–axial hemorrhage sensitivity was 76.0% (CI: 74.0–78.1), similar between IPH (75.4%, CI: 73.5–77.6) and IVH (85.0%, CI: 77.5–92.1). Among parenchymal sites, the model performed significantly better for deep gray nuclei (82.3%, CI: 78.4–86.0) and multilobar hemorrhages (79.2%, CI:75.1–82.9), while worst for occipital lobe (61.4%, CI: 50.0–73.5) and cerebellum (64.8%, CI:56.9–73.2).

Most IVH cases were multi-compartmental (1136 out of 1215), while only 79 were isolated. For cases with IVH only, sensitivity was 85% (CI: 77.5–92.1) with maximum 100% for diffuse IVH (Table S5).

The model performed significantly worse on hemorrhages <10 mm (74.8%, CI:72.8–76.8) compared to large hemorrhages >10 mm (95.0%, CI: 94.1–94.8) and in the absence of mass effect (75.0%, CI: 73.5–76.5) compared to cases with mass effect (88.0%, CI: 86.9–89.0).

Figure 2 summarises model sensitivity by hemorrhage acuity and subgroup in a sunburst plot.

Fig. 2: Model sensitivity across hemorrhage subgroups stratified by acuity.
Fig. 2: Model sensitivity across hemorrhage subgroups stratified by acuity.
Full size image

A Acute intracranial hemorrhage cases. B Non-acute intracranial hemorrhage cases (subacute and chronic). Sunburst plots show model sensitivity across hierarchical hemorrhage categories. The innermost ring represents number of compartment involvement (single vs multi-compartment), the middle ring represents hemorrhage compartment (e.g., extra-axial, intraparenchymal, subarachnoid), and the outer ring represents specific anatomical locations or combinations. Colour scale indicates sensitivity (%) from 70 (red) to 100 (green).

Intersectional demographic subgroup analysis

Table S9S11 shows detailed intersection analysis. In patients aged 40–65 and ≥65 years, sensitivity was significantly lower in the outpatient setting (70.8%, CI: 61.9–78.8 and 73.1%, CI: 68.5–77.4, respectively) compared with other settings, while specificity was significantly higher in the ED (98.2%, CI: 98.1–98.4 and 97.9%, CI: 97.7–98.0, respectively).

Performance analysis based on the intersection of age and race shows that in white patients, sensitivity is significantly lower in younger patients ( < 40 years: 74.9%, 95% CI: 67.7–82.1) compared with older groups (40–65 years: 81.5%, 95% CI: 78.3–84.4; ≥65 years: 84.6%, 95% CI: 82.8–86.3), while specificity is highest in the youngest subgroup (98.3%, 95% CI: 97.9–98.7). In contrast, no statistically significant differences in sensitivity or specificity were observed across age subgroups for Asian patients.

When examining race and patient–class intersections, we found that in both black and white patients, sensitivity is significantly lower in the outpatient setting (black: 72.0%, 95% CI: 64.2–79.2; white: 69.8%, 95% CI: 64.0–75.2) compared with the ED (black: 82.2%, 95% CI: 79.6–84.5; white: 85.4%, 95% CI: 82.7–88.0) and inpatient settings (black: 80.7%, 95% CI: 79.2–82.3; white: 84.4%, 95% CI: 82.4–86.1). In Asian patients, sensitivity was also lowest in the outpatient setting (75.1%, 95% CI: 63.0–86.4), though confidence intervals overlapped with those of the other settings. Across all racial groups, specificity was significantly higher in the ED compared with inpatient or outpatient settings.

Multivariate Analysis

In multivariable analysis, hematoma size, acuity, number of compartments, and mass effect were all independently associated with the model performance (Table S12). Hematomas >10 mm were associated with nearly four–fold higher odds compared with those ≤10 mm (OR 3.82, CI: 3.02–4.85). Acute hemorrhages carried substantially greater risk than subacute/chronic cases (OR 5.93, CI: 4.72–7.46). For each additional compartment involved the odds increased more than threefold (OR 3.25, CI: 2.79–3.79), and the presence of mass effect was associated with a 62% increase in odds (OR 1.62, CI: 1.41–1.87).

Discussion

Our analysis identified an overall sensitivity of 82.2% (CI: 81.3–83.1) and specificity of 97.6% (CI: 97.5–97.7) for ICH diagnosis, compared to FDA clearance metrics of 96.2% (CI: 90.4–98.9) and 94.8% (CI: 89.1–98.1), respectively. These results are consistent with prior external validation studies of the same tool, which have demonstrated wide ranges of sensitivities from 68.2% to 92.3% and specificities between 94.0% and 99.2%, with overall accuracies above 90%13,14,15,16,17,18,19,20.

Model performance varied significantly with hemorrhage acuity, performing best for acute hemorrhages (86.2%, CI: 85.1–87.3) and acute–on–chronic hemorrhages (89.6%, CI: 85.7–93.3), which appear hyperdense, and worst for subacute hemorrhages (45.5%, CI: 38.3–52.0), which are often iso–dense to adjacent parenchyma. Performance for chronic hemorrhages (usually hypodense) was also poor (54.8%, CI: 49.0–60.9). These findings suggest that density is a critical factor for ICH detection models21. Extent of hemorrhage significantly influenced model performance in ICH detection. Detection performance for all single–compartment hemorrhages was significantly lower (76.0%, CI: 74.7–77.2) than any combination multi–compartment cases (93.6%, CI: 92.6–95.5). This is consistent with prior work showing significant variation in model performance between single- and multi-compartment hemorrhages20, and worse performance for subtle subdural and subarachnoid hemorrhages14,15,16,22,23. Performance was also much lower for small hemorrhages <10 mm (sensitivity 74.8%, CI: 72.8–76.8) and hemorrhages in the posterior fossa (sensitivity 61.4%–64.8%).

Previous studies have shown that the model has the lowest sensitivity and PPV in outpatient cases16,22. Our findings are consistent with this pattern, as sensitivity was lowest in outpatients (72.2%, CI: 68.4–76.1) compared with other patient locations, a trend that was also evident across age and racial subgroups. This discrepancy may be explained by differences in the distribution of key imaging features. In our dataset, small hemorrhages (52.1%) and single-compartment hemorrhages (85.9%), two independent features affecting model performance, were significantly more common in outpatients. Additionally, subacute and chronic hemorrhages which are inherently more challenging to diagnose, were more frequently observed in this group (Table S6). Demographic disparity was not well evaluated in prior studies, except one study which demonstrated lower PPV for Black patients and females, and no disparity among age24. Our results show no significant differences in sensitivity or specificity in models’ performance across age, sex, and race subgroups.

Examination of representative examples of false positives predictions and associated heatmaps provided by the software revealed thathyperdense lesions, such as calcifications, vascular malformations, tumors, iodinated contrast, and/or streak artifacts were largely responsible for the erroneous predictions (Figs. S2S4). This is consistent with prior studies which demonstrated false positives from contrast staining in ischemic stroke patients, beam-hardening, post-operative changes, tumors, calcified falx cerebri, artifacts, and postoperative dural thickening25. Among these, calcification and radioiodine contrast are among the main mimickers of ICH26. Howevere, this may not have an impact on real-world utility since these artifacts are easily seen by radiologists. Because the AI software provides a heatmap, the radiologist can see if the region of interest responsible for the prediction is a real ICH or artifact. In contrast, false negatives arose when hemorrhages were either iso-dense (hyperacute or subacute) or hypodense (chronic) compared to surrounding parenchyma, or when artifacts obscure their visualization especially in subarachnoid and subdural hemorrahges. This is in keeping with prior work showing that 63.3% of false negatives were subacute ICH16. Also in one study 39% and 33% of missed ICH cases among residents were SDH and SAH, retrospectively27.

Overall, our results highlight a consistent pattern: the model performs best when hemorrhages are extensive or span multiple compartments, but its sensitivity declines for subtle, localized, or otherwise challenging presentations, particularly in subdural and subarachnoid regions. This is similar to human performance in detection of hemorrhage28,29 and underscore the importance of understanding model bias when deploying such tools in clinical settings.

Our findings support the need to more deeply evaluate model performance in post-deployment, real-world settings. FDA-clearance of AI models often depends on small and opaque datasets, without information on case distribution, severity, demographics, or technical factors leading to underperformance on clinical deployment. We also emphasize the need to consider pathology or imaging findings as more than a binary classification task. Findings of different severity have varying difficulty of detection for physicians30,31, with important, different clinical impact. Therefore, models that underperform on the same types of findings that are difficult to detect for human readers have significantly attenuated clinical utility. Instead, AI models should augment human performance by performing well for situations in which humans underperform. However, in triage settings, such models still offer clinical benefit by expediting hemorrhage identification.

This study has limitations. First, the reference standard was based on final radiology reports authored by board-certified radiologists which is susceptible to the detection performance of the interpreting radiologist. The exams were interpreted with the AI model available, which suggests that radiologists could have both a higher true positive and false positive rates if biased by AI model results. Second, while using an LLM for automated label extraction enabled analysis of a large and richly annotated dataset, it also introduced the potential for mislabeling32. To mitigate this risk, we implemented a structured prompting strategy and verified performance on separate manually annotated holdout sets. Additionally, we provided confidence intervals for all reported performance metrics to reflect underlying uncertainty. Importantly, we validated the model’s subgroup-level performance patterns against a subset of 500 manually labeled cases which confirmed that the relative trends in model behavior across subgroups were consistent with the full cohort, despite some variation in exact performance values. Taken together, these steps support the validity of our findings while acknowledging inherent limitations in reference standard accuracy and data labeling in large-scale retrospective analyses. Residual label noise may persist, particularly in nuanced temporal phrasing, and should be addressed in future hybrid human-LLM labeling frameworks. This study did not quantitatively evaluate time-to-diagnosis or other workflow effects. Future studies should quantify workflow impact while controlling for key confounders, such as case complexity, reader expertise, and protocol/procedure variability. The AI model binary outputs precluded systematic false-positive subclassification; instead, we performed a targeted qualitative review of representative cases and illustrate these in the supplementary material. Although, all non-contrast head CT exams are taken using the same kernels and slice thickness regardless of scanner Scanner vendor and acquisition-parameter subgrouping scanner type could not be analyzed. Finally, this is a single-center study at a large academic institution and tertiary referral center whose patient population may differ than other community settings and therefore affect the real-world experience of the model in those settings.

In conclusion, while the model demonstrated acceptable overall performance, its sensitivity varied significantly by hemorrhage characteristics, clinical setting, and patient demographics. Detection was highest for acute, large, multi-compartment bleeds, but performance declined in single compartment, subacute or small hemorrhages. This affects the model’s role in triaging outpatient cases where most cases are small, localized hemorrhages. These findings highlight the need for targeted model improvements and careful deployment in settings and populations where performance is suboptimal. Future models should be more optimized and evaluated for subtle, small, single-compartment, and non-acute hemorrhages. In addition, transparency of model prediction confidence could help radiologists better understand and trust AI model results. Finally, we emphasize that clinicians and radiologists should treat AI model outputs as decision-support tools, not definitive diagnoses. Over-reliance, particularly in subtle and outpatient presentations, can lead to automation bias and compromised detection of small hemorrhages.

Methods

Patient selection

This was a retrospective, single-center study conducted across a 17-facility academic healthcare system. The study protocol was approved by the Emory University Institutional Review Board (IRB# STUDY00002276) and the research was performed in accordance with the Declaration of Helsinki and relevant institutional guidelines and regulations. The requirement for informed consent was waived by the Emory University Institutional Review Board due to the minimal-risk, retrospective analysis of existing imaging and report data. We included all head NCCT examinations between April 2023 and April 2025 for patients aged ≥18 years that were processed by the Aidoc ICH triage model; studies not processed due to routine technical issues were excluded.

Image acquisition

All examinations were non-contrast head CTs performed as part of clinical care across multiple hospitals, leading to variation in scanner models and manufacturers. Despite this, all sites followed a standardized protocol covering the skull base to vertex without intravenous contrast: Axial thin (0.625 mm) soft tissue kernel, Axial thin (0.625 mm) bone kernel, coronal 2.5 mm soft tissue kernel, sagittal 2.5 mm soft tissue kernel. Per the vendor, the model uses the thinnest available axial slices for inference.

Label extraction using large language model

All CT examinations included in this study were reported using standardized templates by board-certified attending radiologists as part of routine clinical workflow. A ground truth schema of clinically relevant labels was created using consensus from three board certified emergency radiologists, who created definitions and associated keywords for each label. The annotation schema included clinically relevant ICH labels encompassing hemorrhage acuity, anatomical location, post-operative changes, and mass effect (Fig. 3 and Table S1). Importantly, the top-level ICH label contained an option for ‘possible’ to capture when a radiologist expressed uncertainty about presence of a hemorrhage.

Fig. 3: Label extraction hierarchy.
Fig. 3: Label extraction hierarchy.
Full size image

Post-operative state, mass effect, acuity and anatomical compartments were extracted for all ICH positive reports. In cases with multiple hemorrhages, the acuity was selected. Related anatomical location was extracted for each compartment.

We used two distinct evaluation sets for Large Language Model (LLM) performance. Dataset-1 included 500 NCCT radiology reports (350 ICH positive and 150 ICH negative, selected to cover all ICH subtypes) that were manually annotated by radiologists to serve as the primary resource for both developing and evaluating the label extraction prompts. Four hundred of these reports were labeled by radiology residents under the supervision of an attending radiologist (Dataset-1-val). Before annotation, all annotators reviewed a shared expert-defined schema and jointly labeled a small set of sample reports to align on definitions and approach. After this calibration, each report was annotated by a single rater. The remaining 100 reports were separately annotated by three attending board-certified radiologists and discrepancies resolved by consensus and reserved as a holdout validation set (Dataset-1-holdout). Inter-reader agreement is provided in supplementary Table S2.

Separately, 500 randomly selected reports (Dataset-2), were labeled by a board-certified radiologist as ICH positive or negative only to serve as an independent test set to represent the general population prevalence of ICH.

We evaluated the prompt across multiple LLMs and selected GPT-4o (OpenAI, CA, USA) based on highest agreement with manual annotations. Reports were de-identified prior to LLM processing using the Stanford AIMI radiology deidentifier (i2b2-compatible), which removed direct identifiers (e.g., name, MRN, date of birth) from the free-text reports33. We did not alter report content beyond PHI removal to avoid introducing linguistic bias or data leakage. A zero-shot approach was used with no model fine-tuning (temperature = 0). Multiple prompt iterations were systematically tested using Dataset-1-val for refinement; this set was used to adjust prompt wording and extraction logic. The final optimized prompt was selected based on its performance across all labels on the holdout validation set (Dataset-1-holdout). All steps of prompt refining and evaluation were supervised by two board-certified radiologists to ensure medical terminological accuracy and clinical consistency. We explicitly engineered and tested prompt behavior for radiology linguistic edge cases, such as negated acute phrasing (e.g., “no new hemorrhage”) and chronic/stable terminology. Examples of final prompt logic is provided in Supplementary file. The final prompt performance was also evaluated for ICH detection in Dataset-2 as an independent test set.

Labels were extracted hierarchically from the remaining radiology reports (Fig. 3). Reports were classified as ICH positive, negative or possible. Further label extraction was then performed only for ICH-positive cases, including post-operative status, acuity, anatomical compartment of ICH and mass effect. Hemorrhage location and distribution was extracted on a per-compartment basis.

We aggregated certain label classes for the purposes of evaluation. “Multi-compartment” and “diffuse” labels were not extracted directly using a prompt; they were derived from the extracted compartment and location labels. If hemorrhage was present in >1 compartments, the case was assigned as “multi-compartment.” In such cases, size categorization was based on the largest hemorrhage component. For cases with multiple locations within a single anatomical compartment, hemorrhage was classified as “diffuse” for the compartment. In cases with hemorrhages of varying age, the most recent hemorrhage acuity was selected.

AI software predictions

An FDA-cleared commercial deep-learning algorithm (Aidoc Medical, NY) based on a convolutional neural network architecture was used to triage all head NCCT exams for the presence of ICH, providing a binary study-level output (positive or negative). The software operates via a local routing server that forwards studies from PACS to a private cloud for processing. Results are returned to the radiologist’s workstation through a custom widget. Positive results trigger a pop-up with exam details and an optional saliency map (labeled for educational use only). AI model prediction results were received in aggregate from the vendor with case-level information on whether the AI model was executed and its results. Because the AI model runs on all eligible exams, all exams in this study were interpreted with the AI results available to the radiologist (Fig. 4).

Fig. 4
Fig. 4
Full size image

Workflow of AI model inference and comparison with LLM-extracted labels for our evaluation.

Eligible exam selection

Post-operative cases for inpatients were excluded from evaluation due to their complexity, namely the difficulty in distinguishing hemorrhage from post-operative collections, and the limited clinical utility of detecting hemorrhage in this setting. Post-operative cases for outpatients and emergency department (ED) patients were retained, under the assumption that detecting new hemorrhage would be clinically important in these settings. Finally, exams labeled as ‘possible’ hemorrhage were excluded as they could not be confirmed as positive or negative.

Statistical analysis

LLM performance for label extraction was evaluated by comparing its predictions to Dataset-1 and Dataset-2. For binary labels, sensitivity, specificity, and accuracy were computed for each label individually, while for multi-class labels, a confusion matrix was used to calculate accuracy, F1-score, and Cohen’s kappa.

Descriptive statistics and diagnostic performance metrics were used to evaluate the AI model. For overall ICH detection and demographic subgroup analyses, standard classification metrics, including sensitivity, specificity, PPV, NPV, F1-score, and accuracy, were calculated. For imaging subgroups, since the model does not provide subtype predictions, only sensitivity (TP/[TP + FN]) was computed.

To assess the independent contribution of correlated imaging features to model performance, a multivariate logistic regression was performed on ICH-positive cases, with the model’s prediction outcome as the dependent variable. Selected categorical imaging features were included as predictors, and significance was determined at p < 0.05 using maximum likelihood estimation. Subgroup contrasts were pre-specified and limited to clinically defined categories (acute vs subacute vs chronic; ≤10 mm vs >10 mm; single vs multi-compartment). For all metrics and subgroup analyses, 95% confidence intervals (CI) were estimated via non-parametric bootstrap resampling with 1000 iterations. Resampling was performed at the exam level. For subgroup estimates, resampling was stratified by subgroup and ICH status to preserve the observed class mix. CI were taken as the 2.5 and 97.5th percentiles of the bootstrap distribution (fixed random seed).