Longitudinal immune biomarkers and infection risk

Working specification for a pooled analysis in autoimmune disease

pharmacometrics
immunology
references
Can a model of neutrophil, B cell and immunoglobulin time courses, pooled across studies, predict serious infection better than the threshold rules already in use? The specification for the analysis that would answer it.
Published

September 9, 2026

This is a working document, not an analysis. Nothing has been fitted and no dataset has been obtained. Every literature claim below is cited to an entry in References, and every entry there carries ❌, meaning it was characterized from an abstract rather than read. The project index lists the other documents in this folder.

In brief

The question. A patient with an autoimmune disease on B-cell depleting therapy has their neutrophils, lymphocytes and immunoglobulins drawn every few months. Those values are read one at a time against fixed thresholds. Can a model of their time courses, pooled across many studies, predict serious infection better than the thresholds do, and early enough to act on?

Why pooling. Serious infections are rare enough that a single trial produces tens of events, which cannot support a model carrying several biomarkers and their trajectories. The events accumulate across trials, indications and drug classes. Whether those studies can be pooled into one model, or only into a meta-analysis of separate fits, is the first question the work has to answer, and it is answered by looking at data rather than by argument.

The design. A longitudinal submodel for each biomarker, a hazard submodel for time to serious infection, and a link between them that is estimated rather than assumed: current value, slope, or cumulative time under a threshold. Section 7 sets out the three candidate links and Section 9 the comparator each has to beat, which is a rule using the latest value alone.

What the model cannot fix. Prophylaxis is given because a biomarker fell, and dose is held because an infection happened. The measurements and the interventions are entangled, and no amount of pooling separates them without either a randomized comparison or an assumption that cannot be checked from the data. Section 10 states which assumption the analysis makes and what breaks if it is wrong.

How the answer would be used. Three concrete decisions: how often to draw the panel, when to start immunoglobulin replacement or antimicrobial prophylaxis, and whether the next dose is given. A model that improves prediction but changes none of those has not earned its place, and Section 13 says so as a stop rule.


1. High-level objective

Estimate the relationship between the time course of routinely measured immune biomarkers and the hazard of serious infection, using individual patient data pooled across studies, and report whether the trajectory carries information the current value does not.

What comes out of this

  1. A fitted longitudinal model for each biomarker under each drug class, on pooled data, with between-study variability estimated rather than assumed away.
  2. An estimate of which feature of the trajectory enters the infection hazard, with an interval on it.
  3. A comparison against the threshold rules in use today, on the same patients, using time-dependent discrimination and calibration.
  4. A statement of what the data could not answer, with the reason.

Scope

Out of scope, stated once. This is not a mechanistic or quantitative-systems model of immune reconstitution. It is not a vaccine response analysis. It is not a dosing recommendation for any individual drug, and it is not a causal estimate of what prophylaxis does. It predicts, under stated assumptions, and prediction is what it will be judged on.

2. The population

Autoimmune disease, not oncology. The primary population is patients receiving immunosuppressive therapy for an autoimmune indication: rheumatoid arthritis, systemic lupus, ANCA-associated vasculitis, multiple sclerosis, inflammatory bowel disease.

The fork is real, and the reason for the answer belongs here. Oncology has the better data. It has the pooled analyses that already relate a neutrophil time course to an infection event (entries 2 and 11), the myelosuppression model with system parameters that held across drugs (entry 9), and the data sharing platforms. What it does not have is a clean read on the biomarker, because the malignancy itself suppresses the same compartments the drug does, the follow-up is short, and death competes with infection throughout.

Autoimmune disease reverses both. Follow-up runs years, the underlying disease is not itself a haematological malignancy, and infection is the outcome the field already worries about. The oncology literature supplies the methods and the prior estimates; the autoimmune studies supply the patients.

Where an oncology dataset is reachable and an autoimmune one is not, the oncology data is used to develop the machinery and is reported separately, never pooled into the primary fit.

4. The biomarkers to carry

Candidate biomarkers, ordered by how well the trajectory is already understood
Biomarker Arm of the system Behaviour under therapy What practice does with it now
Absolute neutrophil count Innate, bacterial and fungal Falls fast, recovers fast, oscillates with cycles Threshold rules and G-CSF, both long established (entry 1)
CD19+ B cells Humoral, upstream of antibody Falls to below the limit of quantification and stays there for months Watched, seldom acted on
IgG, IgA, IgM Humoral, effector Falls slowly over cycles, recovers over years or not at all Replacement considered in a threshold range (entry 5)
Lymphocytes, CD4+ Cellular, opportunistic Falls with steroids and with several DMARDs Prophylaxis thresholds by convention
CRP, albumin Inflammation and reserve Confounded by disease activity Read as disease activity, not as infection risk

The first three are the primary set. Lymphocyte subsets are carried where reported, which will be a minority of studies. CRP and albumin are carried as covariates rather than as predictors, because in this population they move with the disease as much as with the infection.

The censoring problem is specific to B cells. A depleted count sits at or below the limit of quantification for months, so the trajectory is a long run of values that are not measurements. Handled as left-censored observations in the longitudinal submodel, and it has to be handled explicitly, since substituting half the limit produces a flat line that the link function will happily fit.

5. The infection endpoint

Primary. Time to first serious infection, defined as one requiring intravenous antimicrobials or hospitalization. This definition survives translation across trial and registry sources better than a severity grade does.

Secondary. Time to first infection of any grade; and recurrent events, since a patient who has one serious infection has a raised risk of the next.

Reported separately by organism class where the source allows it: bacterial, viral, opportunistic. This is what Section 3 requires, and it is where the pooled dataset will be thinnest.

Competing risks. Death from other causes, and treatment discontinuation. Discontinuation is not independent of the outcome, since a patient is withdrawn for the thing being predicted.

The endpoint is where the pooling is most likely to fail. Different sources adjudicate differently, count differently, and follow patients for different lengths of time. Section 12 puts a check on this before any model is fitted.

6. The data

What a usable record contains

For each patient: dose and date of every administration of the biologic and of every comedication including glucocorticoids, with dose; every laboratory measurement with its date, its units and its limit of quantification; every infection event with a date, a severity and, where recorded, an organism; the prophylaxis given, with dates; and the baseline covariates in Section 3.

A source that gives biomarker trajectories without dated infection events is useful for the longitudinal submodel and useless for the link.

Candidate sources

Vivli, Project Data Sphere and the YODA Project, described in References. Registry data is a second family: the national biologics registers in rheumatology carry infection events and comedication on a scale no trial has, and carry laboratory values inconsistently.

Harmonization

Units and assays first, then the endpoint. Immunoglobulins are reported in mg/dL and in g/L; cell counts in cells/µL and in 109/L. Flow cytometry gating for B cells differs by laboratory and its lower limit differs by more.

7. The model

7a. Longitudinal submodel

One submodel per biomarker, on the log scale, with drug exposure driving a depletion and a recovery. For neutrophils the structure is settled: the transit-compartment model with system parameters that held across drugs (entry 9) is the starting point. For B cells and immunoglobulins it is not settled, and the first fits should compare an indirect-response form against a flexible spline in time, since a slow recovery over years may not be a feedback process at all.

Between-study variability is a random effect at the study level on the parameters that assay and population differences would move, which is a choice made in the fitting rather than declared here.

7c. Hazard submodel

Parametric baseline hazard, stratified by study or by drug class. Covariates from Section 3, with steroid dose time-varying. Recurrent events as a secondary analysis.

8. Alternatives to the joint model

Stated here so the choice is made deliberately, and settled by entry 16.

Landmarking. At a fixed time, take everything measured so far, fit an ordinary survival model, and predict forward. Simpler, pools across heterogeneous studies more easily, and gives up the measurement-error handling.

Two-stage. Fit the longitudinal model, take the individual predictions, put them into the hazard model as covariates. Standard errors are wrong unless corrected, and the correction is the reason joint estimation exists.

Model-based meta-analysis on aggregate data. The fallback when individual data cannot be obtained. Mean trajectories digitized from publications, event rates from the same publications, and a model relating the two at the study level. It answers a different question, since a between-study relationship is not a within-patient one, and the specification should say that in whatever it reports rather than in a footnote.

The decision. Start with landmarking to establish whether any signal exists, then fit the joint model. Landmarking answers the question that decides whether to continue, in a fraction of the time.

9. The comparator, and what counts as better

The comparator is the rule in use today: the latest value of one biomarker against a fixed threshold. A second comparator is a baseline-only model, since a score computed once before treatment is what has actually reached clinical use in an adjacent field (entry 12).

Reported for each: time-dependent discrimination, calibration over a prediction horizon, and the Brier score. Validation is leave-one-study-out, because the question is whether the model transfers to a study it has not seen, which is the only version of the question that decides whether it can be used.

A model that beats the threshold rule on discrimination but not on calibration has not delivered anything usable, since the decisions in Section 1 are made against an absolute risk.

10. Confounding and the response-to-measurement loop

The dashed arrows in Section 3 are the estimation problem.

What the analysis assumes. That prophylaxis, replacement and dose holds are given in response to observed history, and that this history is recorded. Under that assumption the intervention is a time-varying covariate and the estimated link is interpretable as a prediction, though not as a causal effect.

What breaks it. A clinician acting on something not in the data, which is most of what a clinician acts on. The effect is to bias the biomarker link toward zero, since the highest risk patients get corrected.

What is done about it. Three things. Report the analysis with and without the intervention covariates, since the gap between them measures the problem. Use entry 7, a randomized trial that raised immunoglobulin on purpose, as an external check on the sign and rough size of the effect. And state in the report that the estimate is a prediction under an observed-treatment policy, not an effect of a level.

11. Simulate before requesting data

Requesting individual patient data takes months. Before any of it is requested, simulate the pooled dataset and fit the model to it.

What the simulation has to establish: how many patients and how many events are needed to distinguish the three links in Section 7b from each other; what a realistic measurement schedule, quarterly rather than daily, does to that; what left-censored B cell values do to the estimate; and how much between-study variability the design tolerates before the pooled estimate is uninformative.

If the simulation says the three links are not separable at any plausible number of events, the project stops here, and that is a result to write down.

12. Implementation sequence

  1. Simulation study of Section 11. No real data.
  2. Read the queue in References, in the order given there, and record what each changes in this document.
  3. One dataset, one drug class, one biomarker. Fit the longitudinal submodel and the landmark analysis. This is the point at which the endpoint harmonization problem becomes concrete.
  4. Add the second and third biomarkers. Fit the joint model.
  5. Add the second study. This is the first test of pooling, and the first opportunity to find out that it does not work.
  6. Leave-one-study-out validation against the comparators in Section 9.
  7. Report, including the failures.

13. Stop rules

The project stops, and says why, if any of the following holds.

  • The simulation in Section 11 shows the links are not separable at plausible event counts.
  • No two reachable sources share an infection endpoint that can be made common.
  • The trajectory adds nothing over the latest value on leave-one-study-out validation.
  • The model beats the comparator but changes none of the three decisions in the In brief section.

14. Open questions

  1. Which indication first. Rheumatoid arthritis has the registries and the comedication complexity. Multiple sclerosis has the cleanest anti-CD20 trajectories and the pooled sponsor analysis (entry 4). Undecided.
  2. Whether the biologic is modelled as an exposure or as a group. Carrying pharmacokinetics for each drug is more faithful and multiplies the data requirement.
  3. Whether COVID-19 era events are in or out. They dominate infection counts in any dataset spanning 2020 to 2022, and they change the organism mix entirely.
  4. Paediatric and primary immunodeficiency data. Immunoglobulin trajectories there run for decades against a recorded infection history, and no search of that literature has been run.
Back to top