Human + AI ICU Prediction Laboratory
HUMAN + AI · SYNTHETIC FORECAST EVALUATION
Human + AI ICU Prediction Laboratory
Synthetic AKI within 48 hours of a fictional ICU snapshotSELECTED SYNTHETIC RECORD
Synthetic case S-1594
Realized synthetic outcome AKI did not occur
Ensemble arithmetic 0.70 × 39.9% +
0.30 × 71.0% =
49.2%
The simulator records a binary outcome and evaluated forecasts. It does not claim an observable “true individual risk”.
Cohort Brier scores
Clinician
0.1208Model
0.1412Ensemble
0.1170Gain vs better member
+0.0039Mean forecasts: clinician 20.3% · model 20.4% · ensemble 20.4%. Observed synthetic event rate: 21.1%.
WHAT HAPPENED?
The pool beats both individual forecasts on this synthetic cohort.
The two probability streams contribute useful differences at the current weight, so their convex average has the lowest Brier score of the three.
- The ensemble improves on the better member by 0.0039 Brier-score units in this generated cohort.
- Forecast differences reduce the pool below the weighted-average member loss by 0.0100 Brier-score units.
Why the result happened
PROBABILITY LOSS
Brier score across the mixture
Gain versus better member +0.0039
Clinician Brier
0.1208 AUC 0.836Model Brier
0.1412 AUC 0.758Ensemble Brier
0.1170 AUC 0.869Constant-prevalence reference 0.1665 at an observed synthetic event rate of 21.1%.
Hindsight minimum w = 0.71 · Brier 0.1169.
Lowest score in hindsight on this same synthetic cohort—not a deployable weight.
The pool’s Brier score (0.1170) is no worse than the weighted-average loss (0.1270) by a diversity term of 0.0100. It can still be worse than the better member.
01 How this synthetic cohort works+
The declared endpoint is Synthetic AKI within 48 hours of a fictional ICU snapshot. A fixed seed creates 3,000 wholly synthetic records. Controls transform the same retained base draws, so changing only the ensemble weight cannot change either member’s forecast.
- Seed
icu-lab-v1- Development event rate
- 20.0%
- Deployment event rate
- 20.0%
- Shared residual ρ
- 0.05
1 · Synthetic outcome
Yᵢ ∼ Bernoulli(πdeployment)
In ordinary language: each fictional record receives a realized yes/no outcome using the configured deployment event rate. Conditional score sampling is a joint-distribution device; it does not mean a forecaster “saw the future”.
2 · Correlated hidden residuals
εH,i = ZH,i; εA,i = ρZH,i + √(1 − ρ²) ZA,i
Two independent standard-normal draws are mixed so the hidden clinician and model score residuals have the requested conditional dependence. This is not an observed clinical correlation.
3 · Discrimination and probability
dj = √2 Φ⁻¹(Aj); Sij = (2Yi − 1)dj/2 + εij
The target AUC becomes separation between outcome classes in a binormal score model. The score places synthetic event and non-event records into overlapping normal distributions with the requested population ranking ability, subject to finite-cohort variation.
4 · Development probability and calibration transform
η(0)ij = logit(πdevelopment) + djSij; p(0)ij = logistic(η(0)ij)
The latent score becomes a probability anchored to the development population’s event rate. When development and deployment rates match, this base probability is calibrated in the population.
pij = logistic((logit(p(0)ij) − cj) / mj)
The configured conventional calibration intercept c and positive slope m transform those log odds. Zero intercept and slope one leave the base probability unchanged.
logit P(Y = 1 ∣ pij) = cj + mj logit(pij)
In matched development and deployment populations, this is the conventional calibration relation: zero intercept and unit slope are ideal. A deployment prevalence change breaks that match without applying the shift twice.
5 · Weighted pool
pE,i = w pH,i + (1 − w) pA,i
The ensemble is ordinary convex arithmetic: 0.70 of the clinician probability plus 0.30 of the model probability.
02 What the metrics mean+
Brier score
BS(p) = (1/N) Σᵢ (pᵢ − Yᵢ)²
The mean squared difference between a probability and the realized binary outcome. Lower is better probabilistic performance, but this does not measure treatment utility or clinical usefulness.
Ensemble gain
min(BSH, BSA) − BSE
Positive means the pool beats the better member on this cohort. Current gain: 0.0039.
Cross-error identity
BSE = w²BSH + (1 − w)²BSA + 2w(1 − w)C
C is the mean product of the two signed forecast errors. It makes shared error explicit. Numerical identity residual: -4.16e-17.
Why averaging can help—but not magically
wBSH + (1 − w)BSA − BSE = w(1 − w) mean[(pH − pA)²] ≥ 0
A convex pool is no worse than the weighted average of member losses because disagreement can create a diversity benefit. It can still lose to the better individual member. Current diversity term: 0.0100; identity residual: 7.29e-17.
AUC and calibration
AUC describes ranking, not probability honesty. Calibration compares mean forecast probability with observed event frequency; Wilson intervals expose finite-bin uncertainty. Neither alone determines a treatment threshold.
03 View chart data+
Cohort summary
| Series | Mean prediction | Brier score | Realized AUC |
|---|---|---|---|
| Clinician | 20.3% | 0.1208 | 0.836 |
| Model | 20.4% | 0.1412 | 0.758 |
| Ensemble | 20.4% | 0.1170 | 0.869 |
| Constant prevalence | 21.1% | 0.1665 | — |
633 events among 3,000 synthetic cases.
Weight curve
| Clinician weight | Model weight | Ensemble Brier |
|---|---|---|
| 0.00 | 1.00 | 0.14121 |
| 0.01 | 0.99 | 0.14053 |
| 0.02 | 0.98 | 0.13987 |
| 0.03 | 0.97 | 0.13921 |
| 0.04 | 0.96 | 0.13856 |
| 0.05 | 0.95 | 0.13793 |
| 0.06 | 0.94 | 0.13730 |
| 0.07 | 0.93 | 0.13668 |
| 0.08 | 0.92 | 0.13607 |
| 0.09 | 0.91 | 0.13547 |
| 0.10 | 0.90 | 0.13488 |
| 0.11 | 0.89 | 0.13430 |
| 0.12 | 0.88 | 0.13373 |
| 0.13 | 0.87 | 0.13317 |
| 0.14 | 0.86 | 0.13262 |
| 0.15 | 0.85 | 0.13208 |
| 0.16 | 0.84 | 0.13155 |
| 0.17 | 0.83 | 0.13103 |
| 0.18 | 0.82 | 0.13051 |
| 0.19 | 0.81 | 0.13001 |
| 0.20 | 0.80 | 0.12951 |
| 0.21 | 0.79 | 0.12903 |
| 0.22 | 0.78 | 0.12856 |
| 0.23 | 0.77 | 0.12809 |
| 0.24 | 0.76 | 0.12763 |
| 0.25 | 0.75 | 0.12719 |
| 0.26 | 0.74 | 0.12675 |
| 0.27 | 0.73 | 0.12632 |
| 0.28 | 0.72 | 0.12590 |
| 0.29 | 0.71 | 0.12550 |
| 0.30 | 0.70 | 0.12510 |
| 0.31 | 0.69 | 0.12471 |
| 0.32 | 0.68 | 0.12433 |
| 0.33 | 0.67 | 0.12396 |
| 0.34 | 0.66 | 0.12360 |
| 0.35 | 0.65 | 0.12325 |
| 0.36 | 0.64 | 0.12290 |
| 0.37 | 0.63 | 0.12257 |
| 0.38 | 0.62 | 0.12225 |
| 0.39 | 0.61 | 0.12194 |
| 0.40 | 0.60 | 0.12163 |
| 0.41 | 0.59 | 0.12134 |
| 0.42 | 0.58 | 0.12105 |
| 0.43 | 0.57 | 0.12078 |
| 0.44 | 0.56 | 0.12051 |
| 0.45 | 0.55 | 0.12026 |
| 0.46 | 0.54 | 0.12001 |
| 0.47 | 0.53 | 0.11977 |
| 0.48 | 0.52 | 0.11955 |
| 0.49 | 0.51 | 0.11933 |
| 0.50 | 0.50 | 0.11912 |
| 0.51 | 0.49 | 0.11892 |
| 0.52 | 0.48 | 0.11873 |
| 0.53 | 0.47 | 0.11855 |
| 0.54 | 0.46 | 0.11838 |
| 0.55 | 0.45 | 0.11822 |
| 0.56 | 0.44 | 0.11807 |
| 0.57 | 0.43 | 0.11793 |
| 0.58 | 0.42 | 0.11780 |
| 0.59 | 0.41 | 0.11767 |
| 0.60 | 0.40 | 0.11756 |
| 0.61 | 0.39 | 0.11746 |
| 0.62 | 0.38 | 0.11736 |
| 0.63 | 0.37 | 0.11728 |
| 0.64 | 0.36 | 0.11720 |
| 0.65 | 0.35 | 0.11714 |
| 0.66 | 0.34 | 0.11708 |
| 0.67 | 0.33 | 0.11704 |
| 0.68 | 0.32 | 0.11700 |
| 0.69 | 0.31 | 0.11697 |
| 0.70 | 0.30 | 0.11695 |
| 0.71 | 0.29 | 0.11694 |
| 0.72 | 0.28 | 0.11695 |
| 0.73 | 0.27 | 0.11696 |
| 0.74 | 0.26 | 0.11698 |
| 0.75 | 0.25 | 0.11701 |
| 0.76 | 0.24 | 0.11705 |
| 0.77 | 0.23 | 0.11709 |
| 0.78 | 0.22 | 0.11715 |
| 0.79 | 0.21 | 0.11722 |
| 0.80 | 0.20 | 0.11730 |
| 0.81 | 0.19 | 0.11738 |
| 0.82 | 0.18 | 0.11748 |
| 0.83 | 0.17 | 0.11759 |
| 0.84 | 0.16 | 0.11770 |
| 0.85 | 0.15 | 0.11783 |
| 0.86 | 0.14 | 0.11796 |
| 0.87 | 0.13 | 0.11811 |
| 0.88 | 0.12 | 0.11826 |
| 0.89 | 0.11 | 0.11842 |
| 0.90 | 0.10 | 0.11860 |
| 0.91 | 0.09 | 0.11878 |
| 0.92 | 0.08 | 0.11897 |
| 0.93 | 0.07 | 0.11917 |
| 0.94 | 0.06 | 0.11938 |
| 0.95 | 0.05 | 0.11960 |
| 0.96 | 0.04 | 0.11983 |
| 0.97 | 0.03 | 0.12007 |
| 0.98 | 0.02 | 0.12032 |
| 0.99 | 0.01 | 0.12058 |
| 1.00 | 0.00 | 0.12085 |
Reliability aggregates
| Series | Bin | Cases | Events | Mean forecast | Event frequency | Wilson 95% |
|---|---|---|---|---|---|---|
| Clinician | 1 | 300 | 4 | 1.5% | 1.3% | 0.5%–3.4% |
| Clinician | 2 | 300 | 7 | 3.4% | 2.3% | 1.1%–4.7% |
| Clinician | 3 | 300 | 17 | 5.6% | 5.7% | 3.6%–8.9% |
| Clinician | 4 | 300 | 14 | 8.3% | 4.7% | 2.8%–7.7% |
| Clinician | 5 | 300 | 35 | 11.4% | 11.7% | 8.5%–15.8% |
| Clinician | 6 | 300 | 53 | 15.3% | 17.7% | 13.8%–22.4% |
| Clinician | 7 | 300 | 64 | 20.9% | 21.3% | 17.1%–26.3% |
| Clinician | 8 | 300 | 92 | 29.0% | 30.7% | 25.7%–36.1% |
| Clinician | 9 | 300 | 140 | 42.3% | 46.7% | 41.1%–52.3% |
| Clinician | 10 | 300 | 207 | 65.6% | 69.0% | 63.6%–74.0% |
| Model | 1 | 300 | 12 | 3.4% | 4.0% | 2.3%–6.9% |
| Model | 2 | 300 | 23 | 6.4% | 7.7% | 5.2%–11.2% |
| Model | 3 | 300 | 26 | 8.8% | 8.7% | 6.0%–12.4% |
| Model | 4 | 300 | 34 | 11.5% | 11.3% | 8.2%–15.4% |
| Model | 5 | 300 | 39 | 14.6% | 13.0% | 9.7%–17.3% |
| Model | 6 | 300 | 48 | 18.1% | 16.0% | 12.3%–20.6% |
| Model | 7 | 300 | 71 | 22.6% | 23.7% | 19.2%–28.8% |
| Model | 8 | 300 | 87 | 28.2% | 29.0% | 24.2%–34.4% |
| Model | 9 | 300 | 117 | 36.5% | 39.0% | 33.7%–44.6% |
| Model | 10 | 300 | 176 | 53.9% | 58.7% | 53.0%–64.1% |
| Ensemble | 1 | 300 | 2 | 3.9% | 0.7% | 0.2%–2.4% |
| Ensemble | 2 | 300 | 5 | 6.5% | 1.7% | 0.7%–3.8% |
| Ensemble | 3 | 300 | 6 | 8.9% | 2.0% | 0.9%–4.3% |
| Ensemble | 4 | 300 | 18 | 11.4% | 6.0% | 3.8%–9.3% |
| Ensemble | 5 | 300 | 30 | 14.0% | 10.0% | 7.1%–13.9% |
| Ensemble | 6 | 300 | 38 | 17.1% | 12.7% | 9.4%–16.9% |
| Ensemble | 7 | 300 | 61 | 21.4% | 20.3% | 16.2%–25.2% |
| Ensemble | 8 | 300 | 93 | 27.5% | 31.0% | 26.0%–36.4% |
| Ensemble | 9 | 300 | 151 | 37.5% | 50.3% | 44.7%–56.0% |
| Ensemble | 10 | 300 | 229 | 55.4% | 76.3% | 71.2%–80.8% |
04 What this laboratory leaves out+
This laboratory is a forecast-evaluation simulator, not an AKI model. The simulated “clinician” is a probability-generating process, not a model of human cognition. Version one does not perform:
- Real clinical prediction
- Model training or retraining
- Missing-data imputation
- EHR or FHIR ingestion
- Patient upload or patient-level interpretation
- Decision thresholds, decision curves, or net-benefit analysis
- Treatment recommendations
- Survival analysis, censoring, or competing-risk analysis
- External validation
- Subgroup, equity, or fairness evaluation
- Workflow-impact testing
- Prospective monitoring
- Clinician-behaviour modelling
- General covariate or concept shift
- A causal simulation of hospital artifacts
- Statistical claims about actual clinicians or AI systems
The shared-artifact preset uses correlated hidden residuals as a proxy, not a causal account of how an artifact is learned. The population-shift preset changes prevalence only; real shifts are broader. No value is an estimate of a real clinician, model, hospital, or AKI incidence.
The synthetic laboratory is ready.
Uses the speech voice supplied by your browser or device.
Educational simulation only. Every case, forecast, and outcome in this laboratory is synthetic. It does not estimate any real patient’s risk and must not be used for diagnosis, triage, treatment, monitoring, or any other clinical decision.
Fixed endpoint: synthetic AKI within 48 hours of a fictional ICU snapshot.
1. The arithmetic is easy; the error structure is not
Combining two probability forecasts looks almost insultingly simple. Give the simulated clinician a weight, give the simulated model the remainder, multiply, and add. Yet the result depends less on the elegance of the arithmetic than on what the two forecasts know, what they miss, and whether they fail in the same places.
The useful metaphor is two clocks. Averaging their readings can reduce error when the clocks drift differently. Two clocks wired to the same bad battery may agree beautifully about the wrong time.
That is the purpose of the laboratory above. It does not ask whether “human” or “AI” wins in general. It asks a narrower, inspectable question: on one reproducible synthetic cohort, how do discrimination, calibration, shared residual structure, and ensemble weight determine the Brier score of a convex pool?
The first screen gives the answer before the algebra. It places the simulated clinician, model, and ensemble probabilities on the same 0–100% rail for one plainly labelled synthetic case. It then reports their cohort Brier scores and the ensemble gain against the better member. The case is an example of the arithmetic, not evidence about a patient, and the cohort is a generated teaching object, not an ICU dataset.
2. What this laboratory actually simulates
Version one has one binary outcome and no selector: synthetic AKI within 48 hours of a fictional ICU snapshot. The endpoint label is fixed in the implementation so that a future experiment cannot quietly relabel the same scores as mortality, delirium, or another clinical outcome.
Approximately 3,000 cases are generated from a named deterministic seed. A case contains a binary synthetic outcome, two latent scores, two forecast probabilities, one weighted probability, and the losses needed for evaluation. It contains no name, age, sex, diagnosis, laboratory result, comorbidity, bed number, hospital identifier, or supposedly “true individual risk”.
The simulated clinician is a probability-generating process. It is not a cognitive model, a claim about medical judgement, or a synthetic person. The simulated AI is also only a probability-generating process. There is no model training, language model, neural network, cloud service, or hidden inference API behind it.
The generator samples scores conditional on a generated outcome because that binormal construction makes discrimination controllable and population probabilities derivable. This is a convenient way to define a joint probability distribution. It does not mean that a clinician or model sees the future.
All base uniform and Gaussian draws are retained. Moving a slider transforms the same cases. In particular, changing only the ensemble weight cannot change either individual forecast. A secondary action can generate another deterministic cohort, but it lives with the methods because it is a robustness check, not a way to fish for a more pleasing result.
3. Discrimination: ranking is not probability honesty
For a binary outcome, the area under the receiver operating characteristic curve, or AUC, can be read as a ranking probability: select one synthetic event case and one synthetic non-event case, and ask how often the forecaster assigns the event case the higher score. An AUC of 0.5 has no ranking separation; an AUC of 1 has perfect ranking in the population construction.
The simulator converts target AUC for forecaster into a separation between the two conditional normal score distributions:
In ordinary language, converts an AUC to a standard-normal quantile, and multiplication by gives the class separation required by this equal-variance binormal model. The displayed realised AUC can wander slightly because the browser evaluates a finite cohort rather than the infinite population.
AUC is not accuracy, and it does not inspect the numerical honesty of a 20% or 70% forecast. Any strictly increasing transformation preserves ordering. A forecaster can therefore retain a respectable AUC while every probability is systematically too high, systematically too low, too extreme, or too compressed towards the middle.
This is why the controls call the quantity discrimination/AUC rather than “accuracy”, and why changing the positive calibration slope should not materially change realised AUC apart from finite-sample and numerical effects.
4. Calibration: 20% should mean about 20%
Calibration asks a different question. Among synthetic cases assigned roughly probability, does the generated event occur roughly of the time?
The base score becomes a probability calibrated to the development population through:
The first line combines the development event-rate log odds with the forecaster’s scaled latent score. The second line maps those log odds into a probability between zero and one. When development and deployment event rates match, this construction gives population-calibrated probabilities before the optional calibration controls are applied.
The displayed calibration intercept and slope transform that base probability:
This rearrangement creates the conventional calibration relation in which intercept zero and slope one are ideal. A positive intercept means the forecasts are systematically too low; a negative intercept means they are systematically too high. A slope below one describes probabilities that are too extreme, while a slope above one describes probabilities compressed too timidly towards the middle.
The reliability diagram does not make ten bins merely because ten would look tidy. It begins with ten equal-count bins for each forecast, then collapses tied or effectively constant predictions into the number of distinct bins the data can defend. Each point retains its case count, event count, mean prediction, event frequency, and Wilson 95% interval. The interval is sampling uncertainty around a generated bin frequency, not evidence that the synthetic construction represents a clinical population.
Van Calster and colleagues explain why discrimination and calibration must remain separate: useful ranking does not guarantee reliable risks. Their definitions also anchor the control language used here—intercept zero, slope one, lower slopes indicating forecasts that are too extreme, and higher slopes indicating forecasts that are too moderate.
5. Why differently structured errors can help
Let be the simulated clinician forecast, the simulated model forecast, and the clinician weight. The ensemble is:
In plain language, the ensemble is an ordinary convex average. At it is clinician only; at it is model only; everywhere between, the weights are non-negative and sum to one.
For any forecast , the binary Brier score is:
This takes each probability’s distance from its realised zero-or-one synthetic outcome, squares that distance, and averages across all cases. Lower is better for this probabilistic score; zero would mean every probability matched every outcome perfectly.
The exact cross-error term is:
This number averages the product of the two signed forecast errors case by case. When both errors tend to be large in the same direction, the cross term is large; when their departures differ in useful ways, it can be smaller.
Expanding the square of the ensemble error gives:
Nothing magical is hidden here. The ensemble score is determined completely by the two individual scores, the weight, and the way their casewise signed errors overlap. The implementation verifies this identity numerically on every generated result.
A second identity makes the limited guarantee of averaging precise:
The right-hand side is non-negative because it is a weighted average of squared disagreements. Therefore a convex pool cannot be worse than the weighted average of the two members’ Brier losses. But that benchmark is not the better member. If one forecaster is substantially better, placing too much weight on the weaker one can still make the ensemble worse than using the stronger forecaster alone.
The displayed ensemble gain therefore uses the stricter comparison:
A positive value means the pool beats both members on this synthetic cohort. Zero or a negative value means it does not. It is a descriptive same-cohort result, not a confidence interval, external validation, or deployable weighting rule.
6. Why a shared hospital artifact defeats the average
The two hidden residuals begin with independent standard-normal draws and . The simulator imposes conditional dependence through:
Here is the configured correlation between latent residuals after conditioning on the synthetic outcome. At a high positive value, much of the model residual is inherited from the same draw that drives the clinician residual. At lower or negative values, the residuals have more independent or opposing movement.
The preset Both learn the same hospital artifact uses as a proxy for a shared blind spot. It does not causally simulate an artifact, show how one was learned, or establish that a particular EHR field affects real clinicians and models. The title is a thought experiment attached to highly correlated synthetic residuals.
The Shared error view keeps four quantities separate: configured latent residual dependence, realised latent-residual correlation, correlation of casewise squared losses, and the cross-error term . They are not interchangeable. In a binary forecast, the signed errors are mechanically negative when and positive when , so their ordinary correlation is not a clean estimate of the hidden conditional dependence controlled by the slider.
The Canvas plots clinician squared loss horizontally and model squared loss vertically for a deterministic sample of at most 600 cases. Points near the upper right are cases on which both forecasts perform badly; points on opposite sides of the equal-loss diagonal show who was closer for that case. Metrics still use the complete cohort, and a semantic table plus Previous and Next case controls carries the same information without requiring sight or pointer input.
7. Why deployment shift can preserve AUC while damaging calibration
The deployment-shift preset changes the synthetic event rate from 20% in development to 35% in deployment while leaving both explicit calibration controls at intercept zero and slope one. The outcome is drawn at the deployment rate, but the score-to-probability conversion still uses the development rate.
That is a deliberately narrow prior or prevalence shift. The ranking can remain broadly similar because the conditional score distributions retain their separation. The mean predicted probability can nevertheless miss the new observed event frequency, pulling the reliability curves away from the diagonal. Averaging two forecasts anchored to the same obsolete base rate does not invent the missing correction.
The simulator does not also add a prevalence-shift calibration intercept. Doing so would count the same shift twice. Nor does it imply that all deployment shift is prevalence shift. Real systems may face changing covariates, measurement processes, conditional relationships, treatments, documentation, referral patterns, workflows, or outcomes.
This is why external evaluation belongs to a specified population, place, and time. The BMJ guide to evaluating clinical prediction models warns against treating the word “validated” as permanent approval: performance can vary across populations, settings, and periods.
8. What the Brier score rewards—and what it cannot tell us
The Brier score rewards probabilities that stay close to their realised binary outcomes across the cohort. It is an overall probabilistic-performance measure influenced by both discrimination and calibration. The laboratory also reports the score of a constant forecast equal to the observed synthetic prevalence, giving the active cohort a simple no-ranking reference.
The score descends from Glenn W. Brier’s 1950 paper on the verification of probability forecasts. Steyerberg and colleagues’ evaluation framework places it among overall performance measures while keeping discrimination and calibration visible separately.
A lower score does not mean “clinically useful”. Squared error contains no treatment effect, harm, workload, capacity constraint, preference, threshold, or action. This version deliberately omits decision curves and net benefit. A forecast could score better and still fail to improve any decision or outcome. Conversely, a tiny score difference need not matter to any plausible action.
The curve of ensemble Brier score against weight answers one retrospective numerical question on the same synthetic cases. Its marked minimum is labelled as the lowest score in hindsight, not as a deployable weight. Choosing and validating an operational weight would require data and procedures that this laboratory does not contain.
9. Four experiments to try
The preset numbers are illustrative, not empirical estimates:
| Preset | Development rate | Deployment rate | Clinician AUC | Model AUC | Shared residual | Clinician weight |
|---|---|---|---|---|---|---|
| Specialist with missing structured data | 0.20 | 0.20 | 0.82 | 0.75 | 0.05 | 0.70 |
| Trainee plus strong model | 0.20 | 0.20 | 0.70 | 0.86 | 0.15 | 0.15 |
| Both learn the same hospital artifact | 0.20 | 0.20 | 0.83 | 0.83 | 0.95 | 0.50 |
| Deployment population shift | 0.20 | 0.35 | 0.78 | 0.84 | 0.65 | 0.50 |
Both forecasters begin with calibration intercept zero and slope one in every preset. Once a slider moves, the label becomes Custom — based on [preset name]. Reset restores the entire originating preset rather than only the last control.
Start with Specialist with missing structured data. Inspect the gain, then move the clinician weight through the score curve. The reduced model discrimination is the entire representation of “missing structured data”; this is not a simulation of missing at random, missing not at random, extraction failure, measurement delay, or an actual missing-data mechanism.
Choose Trainee plus strong model. Increase clinician weight while watching the two individual scores remain fixed. The experiment exposes the point at which additional weight on a weaker member dilutes the stronger forecast. Then lower shared dependence: complementarity may help, but it cannot guarantee that every weight beats the model alone.
Choose Both learn the same hospital artifact. Record the high latent dependence, casewise loss overlap, and ensemble gain. Now lower while holding both AUC controls fixed. Any improvement comes from changed error structure inside the simulator, not from acquiring new clinical information or modelling the causal history of an artifact.
Choose Deployment population shift. Compare realised AUC with mean predicted and observed event rates, then open Calibration. The forecasts can continue to rank synthetic cases while systematically missing the new rate. Pooling two miscalibrated members may preserve their shared miscalibration or perform worse than the better member.
Only after trying those comparisons should “Generate another synthetic cohort” be used. A new deterministic seed checks whether a qualitative lesson survives another finite draw. It is not permission to repeat cohorts until a preferred conclusion appears.
10. What this laboratory deliberately leaves out
Version one does not perform:
- Real clinical prediction.
- Model training or retraining.
- Missing-data imputation.
- EHR or FHIR ingestion.
- Patient upload or patient-level interpretation.
- Decision thresholds, decision curves, or net-benefit analysis.
- Treatment recommendations.
- Survival analysis, censoring, or competing-risk analysis.
- External validation.
- Subgroup, equity, or fairness evaluation.
- Workflow-impact testing.
- Prospective monitoring.
- Clinician-behaviour modelling.
- General covariate or concept shift.
- A causal simulation of hospital artifacts.
- Statistical claims about actual clinicians or AI systems.
It also omits physiology, longitudinal measurements, competing outcomes, treatment changes, clinician learning, alert effects, selective documentation, measurement error, site clustering, spectrum effects, missingness, dataset shift beyond the single prevalence example, uncertainty around model development, and uncertainty in model selection.
The omission of thresholds is especially important. This is a forecast-evaluation simulator, not an AKI model and not a decision-support trial. Neither a Brier score nor an AUC can establish safety, benefit, fairness, usability, or clinical value without the designs and evidence appropriate to those questions.
11. Methods, reproducibility, and references
Cohort generation
For each synthetic case , the outcome is drawn at the configured deployment rate:
In ordinary language, the case receives outcome one with the deployment probability and outcome zero otherwise, using a deterministic pseudorandom uniform draw. It is a generated label, not an observed event or latent biological truth.
After the correlated residuals and AUC-derived separation have been constructed, each forecaster receives the latent score:
When the generated outcome is one, the class centre is ; when it is zero, the centre is . Adding the residual creates overlapping score distributions whose population AUC is the configured target.
The score becomes development-calibrated log odds:
This line adds evidence from the latent score to the development population’s baseline log odds. Applying the logistic function yields the base probability used by the calibration transformation described earlier.
The inverse-normal calculation uses a stable local approximation tested against known quantiles. Configured is clamped inside its numerical domain before evaluating . Probabilities are clipped only when a logarithm requires protection from exactly zero or one; the displayed probability is not otherwise rounded into a different scientific value.
The named default seed is icu-lab-v1. Uniform and Gaussian base draws are retained so controls transform one cohort instead of silently drawing another. Seed changes are explicit. No default result depends on the current clock, browser entropy, or Math.random().
Evaluation and display
All cohort metrics use the complete set of approximately 3,000 cases. Realised AUC is computed from the generated forecasts. Brier scores, mean probabilities, observed event rate, gain, cross-error term, latent residual correlation, and casewise squared-loss correlation all come from the same immutable evaluated cohort.
Calibration uses equal-count bins separately for each series, collapsing ties when fewer distinct bins are defensible. Wilson intervals expose finite-bin variability. The distribution rug shows where forecasts actually occur, preventing a smooth-looking line from implying support in empty probability regions.
The shared-error Canvas draws no more than a deterministic 600-case sample for rendering performance, and its accessible table mirrors that same display sample. All summary calculations still use the complete cohort. The selected case is a deterministic, moderately high-disagreement record rather than the most extreme outlier, and Previous and Next controls make inspection independent of pointer input.
The score curve evaluates weights from zero to one against the same individual forecasts. Endpoint labels say Model only and Clinician only. Individual-score references and a non-exaggerating scale remain visible. If the member predictions are effectively identical, an “optimal” weight is not identifiable and the display reports an em dash rather than manufacturing certainty.
Everything runs locally. There are no uploads, network inference calls, runtime clinical datasets, analytics transmissions from the instrument, or saved patient-like records.
Evidence boundaries
These sources constrain different claims; none validates the synthetic preset values or turns the laboratory into a clinical model:
- Glenn W. Brier, “Verification of Forecasts Expressed in Terms of Probability”, Monthly Weather Review (1950), is the original probability-score source.
- Steyerberg et al., “Assessing the performance of prediction models: a framework for some traditional and novel measures”, Epidemiology (2010), separates overall performance, discrimination, calibration, and decision-analytic evaluation.
- Van Calster et al., “Calibration: the Achilles heel of predictive analytics”, BMC Medicine (2019), anchors the calibration intercept, slope, curve, and deployment-setting cautions.
- Collins et al., “Evaluation of clinical prediction models (part 1): from development to external validation”, BMJ (2024), explains why performance must be evaluated in intended populations and why a validation study is not permanent approval.
- Collins et al., “TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods”, BMJ (2024), provides harmonised reporting guidance across regression and machine-learning prediction studies.
TRIPOD+AI asks real prediction studies to report their target population, outcome, data sources, model methods, evaluation measures, and limitations transparently. This page borrows that discipline while stating the decisive difference: it is an educational simulation with no clinical dataset, fitted prediction model, validation population, or claim of use.
The useful conclusion is therefore deliberately modest:
An ensemble helps when its members contribute complementary information or differently structured errors. Averaging is arithmetic, not magic. Two forecasters sharing the same blind spot can remain confidently wrong together.
Quick reference and FAQ
Averaging a clinician forecast with an AI-model forecast can help when their errors contain useful differences. It cannot remove a blind spot they share, and it can dilute the better forecast. This wholly synthetic laboratory makes those distinctions visible without estimating any real patient’s risk.
Key Terms
- Probability forecast
- Convex ensemble
- Brier score
- Discrimination
- Area under the curve
- Calibration
- Calibration intercept
- Calibration slope
- Reliability diagram
- Shared residual dependence
- Cross-error term
- Prevalence shift
- Synthetic cohort
Frequently asked questions
Can this laboratory estimate a real ICU patient’s risk of acute kidney injury?
No. Every case, outcome, and forecast is generated by a declared synthetic probability model. The laboratory accepts no patient data and must not be used for diagnosis, triage, treatment, monitoring, or any other clinical decision.
Why can an average be worse than the better individual forecast?
A convex average is no worse than the weighted average of its members’ Brier losses, but that weighted-average benchmark includes some loss from the weaker member. Giving the weaker forecast too much weight can therefore make the ensemble worse than the stronger member alone.
Does a higher AUC mean the forecast probabilities are honest?
No. AUC measures ranking: how often a synthetic event case receives a higher score than a synthetic non-event case. Probabilities can preserve that ranking while being systematically too high, too low, too extreme, or too timid.
What does shared residual dependence mean here?
It is the configured conditional correlation between two hidden Gaussian noise terms in this simulator after conditioning on the synthetic outcome. It is not an observed clinical quantity, the correlation of probability residuals, or proof that a real hospital artifact caused an error.
What does the deployment population shift preset simulate?
Only a prior or prevalence shift: the generated event rate changes from 20% in development to 35% in deployment while the conditional score construction stays fixed. Real population shift can also involve changing covariates, measurements, relationships, workflows, and treatment patterns.
Does a lower Brier score establish clinical usefulness?
No. The Brier score evaluates probabilistic performance by averaging squared forecast errors. It does not encode treatment consequences, decision thresholds, net benefit, workflow effects, safety, or patient outcomes.
Are the displayed differences estimates of real clinicians, AI systems, hospitals, or AKI incidence?
No. The preset values are illustrative controls in a deterministic teaching experiment. They are not empirical estimates, external validations, or claims about any person, organisation, model, or clinical population.
Word Cloud
Related reading
Selected by shared topics and section, with closer publication dates breaking ties.


