Skip to content
Hurdle
Article · Biomarker Science

One Configurable Engine for Risk and Prognostic Biomarkers

Aurora discovers who will develop a disease and who will die from it, from the same cohort engine. Validated on UK Biobank (N ≈ 500,000).


Aurora cohort engine: many candidate signals converge into one engine node, branching into two cohort outputs

Two questions, one engine

Flow diagram: UK Biobank cohort JSON feeds per-period staged filtering and the can_use_in_cohort audit flag, branching into a risk cohort and a prognostic cohort that share a penalised Cox model

The same cohort engine and penalised-Cox backbone, configured for either a risk or a prognostic question.

Risk discovery asks who will develop a disease: cases are the first matching ICD outcome after recruitment. Prognostic discovery asks who will die from a disease they already have: cases are assigned from cause of death. Both run on the same filtering engine and the same time-to-event model, so results stay directly comparable.

Risk biomarkersPrognostic biomarkers
PopulationFiltered by condition codesPatients with the index disease (e.g. I50.* HF)
CasesFirst ICD outcome after recruitmentDeath from a configured cause
Outcome fieldicd_outcome_codesdeath_cause_codes

Table 1 — the same engine configured for risk vs prognostic discovery.

Configurable, reproducible, auditable cohorts

  • Per-period filtering: inclusion and exclusion are evaluated against named time windows, so a comorbidity can be required before recruitment yet excluded before the outcome.
  • Audit by design: the can_use_in_cohort flag records why each participant is kept or dropped, giving transparent filtering statistics for every run.
  • Sensitivity built in: stricter or looser cohort variants (for example a full comorbidity-exclusion block) are one config edit away, with no code changes.

Supported period keys: before_recruitment, all_period, all_period_before_outcome, and after_recruitment_before_outcome (where outcome_end = min(censor_date, date_of_death)).

Proof point: validated on a real cohort

Bar chart: C-index for 15-year CV-death prognosis rises from 0.60 (age, sex baseline) through 0.74 with proteomics to 0.76 for clinical + proteomics + PRS

C-index for 15-year CV-death prognosis among heart failure patients, by data modality. UK Biobank.

On a heart failure to cardiovascular-death cohort (2,870 CV deaths, 13,506 controls), the engine reached a best multimodal C-index of 0.76, with a multi-protein hazard ratio of 3.50 per SD and clinically coherent top features (NT-proBNP, renin, fibrosis and remodelling proteins). The framework is statistically useful and clinically interpretable on the same run.

Generalises to any disease

Pointing Aurora at a new indication means editing the cohort JSON and enabling prognostic mode in the pipeline config. The same engine then supports trial enrichment, patient stratification, and companion-diagnostic development across therapeutic areas, with auditable cohorts and comparable risk and prognostic readouts throughout.

Read nextA Single-Platform Companion Diagnostic for Heart FailureRead the article →
Tom Stubbs, PhD

CEO, Hurdle

He/Him. Tom is CEO at Hurdle, a diagnostic-as-a-service company. Tom is a specialist in Epigenetics, Machine Learning, and Computational Biology.

LinkedIn →

Contact us to scope risk and prognostic biomarker discovery for your therapeutic area.

Talk to us →