PK-platelet modelling to select a Phase 2 dose in ITP

Working specification for a B-cell depleting agent with escalation and expansion data

pharmacometrics
immunology
dose selection
working document
Can a model of pharmacokinetic and platelet-count data for a B-cell depleting agent, fit to a small dose escalation and expansion study, select a Phase 2 dose in immune thrombocytopenia? This specifies the model, the data it needs, and the simulation study that decides whether the answer is yes.
Published

September 9, 2026

Nothing here has been fit, and six sources have been read. The literature review in references.qmd was assembled from search results and abstracts on 2026-09-09. The IWG 2009 consensus paper behind Section 3, the FDA, EMA and PMDA reviews of rilzabrutinib, fostamatinib and efgartigimod behind Section 19, and the eltrombopag ITP model poster behind Sections 8.3 and 8.4 were read in full and are ✅; every other source is ❌. The model in Section 8 is a candidate structure written to be replaced by whatever the rituximab and thrombopoietin receptor agonist papers actually report. The numbers quoted from trials are transcribed from search summaries and are not safe to reuse. The project index lists the other documents in this folder.

In brief

The question. A B-cell depleting agent has been given to a small number of patients with immune thrombocytopenia (ITP) in a dose escalation followed by two expansion arms of 20 to 40 patients each. Serum drug concentration, circulating B-cell count, serum IgG and platelet count were measured serially in every patient. Can a model of those time courses pick the Phase 2 dose, and how many patients does it save over picking from the response rate in each arm?

The answer this specification argues for, before it is tested. ⚠️ Two arms of 20 to 40 can pick the better dose without a model when the doses are 15 or more points apart in durable response rate; at 10 points apart the model is what makes 20 per arm enough; the bottom of a plateau is beyond both. Picking the better of two arms 15 points apart takes 20 per arm and 10 points apart takes 40 per arm on the raw response rates alone. A model that fits the escalation cohorts and both expansion arms through one dose-response shape estimates the same contrast with 2 to 5 times the information, if the shape is right, which brings the 10-point case inside 20 per arm. Where adjacent doses differ by 5 points, as at the bottom of a plateau, no comparison of arms at this size can tell them apart, and the model can name the plateau only as a parameter of a shape it has assumed. Section 14 computes those numbers and Section 15 is the simulation that checks whether the shape assumption survives being wrong. The dose selection is made on predicted platelet response, never on the B-cell biomarker.

What the biomarker in the middle can and cannot say. Circulating B cells are depleted essentially completely at every dose that gets past the first escalation cohort. Depth of depletion is therefore uninformative about dose, and the quantity that still varies with dose is the duration of depletion, measured as time to repopulation. Section 7 is that argument and it is the reason the model in Section 8 is written around a duration. Duration of depletion carries dose into the model; it does not choose the dose, because complete and durable depletion of circulating B cells is not the same thing as maximal platelet response, and the selection rule in Section 13 is written on the platelet endpoint for that reason.

The chain the model has to cross.

flowchart LR
    D["Dose"] --> C["Concentration"]
    C --> B["Circulating B cells<br/>saturates on depth"]
    C --> T["Tissue B cells<br/>spleen, marrow"]
    B --> PB["Plasmablasts"]
    T --> PB
    PB --> A["Anti-platelet IgG<br/>not measured"]
    LLPC["Long-lived plasma cells<br/>not depleted"] --> A
    A --> K["Platelet destruction rate"]
    K --> P["Platelet count"]

Two states in that chain are unobserved in any realistic dataset: tissue B cells and anti-platelet immunoglobulin G (IgG). Both sit between the drug and the endpoint. The model collapses them into one transduction delay whose prior comes from published time-to-response in rituximab-treated ITP, and estimates only the dose-dependent part.

How the question gets answered. By simulation, in Section 15. Generate escalation and expansion datasets from an assumed true model at a realistic sample size, fit the candidate model, apply the dose selection rule in Section 13, and count how often the selected dose is the true best dose. The deliverable is that operating characteristic, per sample size and per dosing scheme, not a dose recommendation for any real compound.

Out of scope, stated once.

  • A recommendation for any named drug. The compound in this specification is generic. Ianalumab, rituximab and mezagitamab appear as sources of realistic parameter values and as external checks, and nothing here is a dosing statement about any of them.
  • A safety model. Dose-limiting toxicity, infection risk and hypogammaglobulinaemia do not enter. The selection rule in Section 13 assumes a safe dose range has already been established and chooses inside it.
  • A mechanistic B-cell systems model. No germinal centres, no B-cell receptor signalling, no compartment-by-compartment trafficking. Published quantitative systems pharmacology models of B-cell differentiation exist and are listed in the references; they are a source of structure, not something to be rebuilt at this sample size.
  • Bleeding. Every IWG response definition carries a bleeding clause (Section 3.1) and this project models platelet count only, so its predicted response rates are upper bounds.
  • Patient-level reanalysis of any published trial. No such data is available here. Public aggregate results are used only as external checks.

1. Objective and preliminary conclusion

1.1 Objective

Determine whether a joint model of drug concentration, circulating B-cell count and platelet count, fit to a Phase 1 dose escalation plus expansion study in primary ITP, supports a Phase 2 dose selection that a later randomized dose comparison would confirm.

The setting assumed throughout:

  • a monoclonal antibody or antibody-like agent whose mechanism is B-cell depletion, given intravenously or subcutaneously as a short course of two to four administrations;
  • an escalation of three to five dose levels, three to six patients per level, whose data are kept and fit alongside the expansion;
  • expansion in two arms at doses chosen from the escalation, twenty to forty patients per arm, so that the whole study is fifty to one hundred patients;
  • serial pharmacokinetic (PK) sampling, circulating B-cell counts by flow cytometry, serum total immunoglobulin G (IgG), and platelet counts at every visit for six to twelve months; anti-platelet IgG where an assay is run;
  • primary clinical endpoint a durable platelet response, defined as platelet count \(\ge 50 \times 10^9\)/L sustained over a specified fraction of the observation window without rescue therapy.

1.2 Preliminary conclusion

The working view, before any of Section 15 has been run, is the statement below, and then an assessment of it.

PK-platelet modelling has not been helpful for agents that prevent platelet destruction, as opposed to agents that enhance production, because of the nature of the endpoint: the endpoint is a dichotomy, and what platelets do in the presence of a B-cell depleter is return to normal. A responder’s count does not become more normal at a higher dose.

⚠️ Correct as a description of the record, correct about the main reason, and it understates what a model can still do. Taking the three parts in turn.

“Has not been helpful”: true of every case found. Three regulatory programmes for destruction-blocking drugs attempted a platelet model (Section 19.1, table). Fostamatinib’s continuous models all failed; efgartigimod’s Phase 2 IgG-to-platelet model could not explain the between-patient variation and the Phase 3 dose was chosen on IgG; rilzabrutinib’s model fit but was built after the dose was chosen and found the durable-response endpoint flat across exposure. None selected a dose. The production drugs, by contrast, have published models that were used for titration (Section 19.2, reason 4).

“Because the endpoint is a dichotomy and platelets return to normal”: the main reason, with two qualifications. The first is that “return to normal” is a simplification. Under the IWG criteria (Section 3.1) a rituximab responder reaches complete response, above 100, in roughly 40% of treated patients and only response, 30 to 100 with a doubling, in another 20 or so; there is a graded layer inside the responders, and the rilzabrutinib review saw it, as a continuous count that rose with exposure while the dichotomy did not. The graded layer is small next to the responder split, and dose acts mostly on the split, but it is not zero. The second is that the dichotomy is not the only reason. Titration to the platelet count and intrapatient escalation each broke exposure-response for at least one of the three programmes (Section 19.2, reasons 5 and 6), and a fixed-dose parallel-arm study removes both without touching the endpoint.

What it understates. If dose acts on the probability of response and not on its size, the model to build is a model of that probability, with the longitudinal count used to classify each patient’s response earlier and more reliably than a threshold on noisy counts does, and the escalation cohorts pooled through a dose-response shape (Sections 8.4, 12 and 14). That model has a bounded gain: it cannot beat one coin flip per patient, and Section 14 puts the bound at a factor of 2 to 5 in patients when the shape is right and nothing when it is wrong. It also answers questions the dichotomy cannot, time to response, durability, and the consequence of a dose interruption or a drug interaction, which is what the rilzabrutinib review used its model for.

So the preliminary conclusion is narrower than the statement. A PK-platelet model of the classical, graded kind will not help here, and the record says so. A dose-to-responder-probability model with a longitudinal component may help when the two candidate doses are within 10 points of each other in response rate, and cannot help when they are further apart, because the raw rates already decide. Whether it helps in the close case depends on a shape assumption that Section 15 exists to test. Run the simulation study; do not yet fit a model to real data.

2. Uncertainties, ranked

Ranked by how much each would change the answer, with where the document deals with it. The first is the one that could make the whole approach wrong, and nothing in the study measures it.

  1. Tissue depletion is not measured, and tissue is where the disease is. The autoreactive plasmablasts and plasma cells that make anti-platelet IgG sit in the spleen and the bone marrow. No ITP trial biopsies either, and a peripheral B-cell count says how well the drug clears blood, not tissue. Ianalumab’s own Phase 2b dose selection in Sjögren’s needed a hypothesized tissue receptor-occupancy model precisely because the circulating count could not stand in for it (references, entry 2). The specification’s response to this is a set of proxies (Section 4: circulating plasmablasts, BAFF, the IgG trajectory, the autoantibody titre where measured), a structural term for the fraction of autoantibody production the drug cannot reach (Section 8.3, \(\phi_i\)), and a wrong-structure case in the simulation where tissue depletion lags peripheral depletion (Section 15). If the wrong-structure case fails, the answer to the project is no, and Section 17 says so first.
  2. The response is a mixture, and the non-responder fraction may depend on dose or may not. Rituximab’s 60% initial and 20 to 30% five-year response rates put the long-lived plasma-cell floor at roughly half the patients. Whether a higher dose or longer depletion lowers that floor is the dose-response the study is trying to see, and it is not known. Sections 8.4 and 12.
  3. The transduction delay is long and wide. Rituximab reaches peak platelet response 14 to 180 days after dosing (Section 3.5). A thirteen-fold spread across patients means a slow responder and a non-responder look the same for months, and the prior on \(k_{e0}\) is doing real work. Section 10.
  4. The dose-response shape that lets the model pool arms is assumed. The factor of 2 to 5 in Section 14 comes from a two-parameter logistic in log dose; a plateau shape pools less and a wrong shape pools confidently wrong. Section 15’s simulation is the check.
  5. The duration signal per patient is weak. One PK half-life of extra depletion per doubling of dose, against a between-patient spread in recovery rate of 22 to 126% CV in published models. Section 6.2.
  6. Concomitant therapy may dominate the platelet trajectory. Most patients entering the study are on a TPO receptor agonist or corticosteroids or both, and rescue therapy is informative censoring. Sections 9 and 12.
  7. The anti-platelet antibody assay is positive in only about half of patients and is semi-quantitative. Where negative, the autoantibody layer stays latent. Sections 4 and 8.5.
  8. Bleeding is in every response definition and not in the model. Predicted response rates are upper bounds. Section 3.1.
  9. The endpoint itself varies by trial in threshold, durability window and rescue rules (Section 3.4), so a response rate is not one number across programmes and the external checks in Section 16 compare shapes rather than rates.

3. Endpoints in ITP

The standard definitions come from the International Working Group (IWG) consensus, Rodeghiero et al., Blood 2009;113:2386-2393. That paper has been read, and the definitions in this section are quoted from it rather than from a summary of it, so they are the only ✅ entries in references.qmd.

There is no partial response in ITP. The IWG panel considered “partial” and “minimal” response, which appear widely in the older literature, and deliberately dropped both, “because of the wide heterogeneity in the criteria used in these definitions.” What replaced them is a two-level scale: complete response and response. A trial reporting a PR in ITP is using a non-standard definition and has to say what it is.

3.1 The IWG response criteria

IWG 2009 response criteria, Table 2 of Rodeghiero et al. ✅
Term Definition
CR, complete response Platelet count \(\ge 100 \times 10^9\)/L and absence of bleeding
R, response Platelet count \(\ge 30 \times 10^9\)/L and at least a 2-fold increase over the baseline count and absence of bleeding
NR, no response Platelet count \(< 30 \times 10^9\)/L or less than a 2-fold increase over baseline or bleeding
Loss of CR Platelet count below \(100 \times 10^9\)/L or bleeding
Loss of R Platelet count below \(30 \times 10^9\)/L or less than a 2-fold increase over baseline or bleeding
Time to response Time from starting treatment to achievement of CR or R
Duration of response From achievement of CR or R to loss of CR or R; also reported as the proportion of cumulative time spent in CR or R over the period examined
Corticosteroid dependence Ongoing or repeated corticosteroids for at least 2 months to hold platelets at or above \(30 \times 10^9\)/L or to avoid bleeding. Corticosteroid-dependent patients count as non-responders

Three qualifications from the same table, each of which changes what a model has to simulate.

A response is a confirmed response. Platelet counts must be confirmed on at least two separate occasions, at least 7 days apart when used to define CR or R, and 1 day apart when used to define NR or loss of response. A single count above threshold is not an event, which is why Section 15 simulates the visit schedule rather than the underlying trajectory.

Bleeding is part of every definition. CR, R and NR each carry a bleeding clause. A model of platelet count alone predicts a necessary condition for response, not response. The specification treats bleeding as out of scope and this is where that costs something: predicted response rates from Section 13 are upper bounds.

Baseline is the count at the start of the investigated treatment. The 2-fold increase in the definition of R is relative to that count, so baseline severity enters the endpoint itself and not only the model. This is the second argument for requirement 6 in Section 9.

3.2 Disease phases and severity

IWG 2009 disease phases, Tables 1 and 4 ✅
Term Definition
Primary ITP Isolated thrombocytopenia, platelet count \(< 100 \times 10^9\)/L, with no other cause
Newly diagnosed Within 3 months of diagnosis
Persistent Between 3 and 12 months from diagnosis
Chronic Lasting more than 12 months
Severe Bleeding at presentation sufficient to mandate treatment, or new bleeding requiring a different platelet-enhancing agent or an increased dose
Refractory Failure to achieve at least R, or loss of R, after splenectomy, plus a need for treatment to minimize bleeding risk

The phase decides the trial population and therefore the response rate the model has to reproduce. Persistent and chronic ITP is where every agent in Section 16’s external checks was studied, and spontaneous remission is common enough in newly diagnosed ITP that a placebo arm there behaves differently.

3.3 Primary and secondary endpoints as the IWG specifies them for trials

Table 5 of the same paper gives trial-adapted criteria.

  • Primary endpoint: CR or R based on platelet count, as in Section 3.1.
  • Secondary endpoints: adverse events; need for rescue interventions; corticosteroid and concomitant treatment reduction; rate of splenectomy; bleeding measured on a validated scale; health-related quality of life on a validated instrument; and pharmacoeconomic analysis where possible. The paper notes that any of these may become the primary endpoint depending on the trial design.
  • Eligibility: entry platelet count below \(30 \times 10^9\)/L, or below \(50 \times 10^9\)/L in specific settings, on steroids, or with bleeding.
  • A predefined exceedingly high platelet count induced by treatment may be recorded as an adverse event. A model whose drug effect is unbounded above has to simulate that too.

3.4 What recent trials actually use

Trials have drifted away from the IWG thresholds, toward \(50 \times 10^9\)/L with a durability requirement. Every definition below is transcribed from a search summary and carries ❌ in the references.

Endpoint definitions in recent ITP trials ❌
Trial, drug Primary endpoint as reported
LUNA3, rilzabrutinib Platelet \(\ge 50 \times 10^9\)/L for \(\ge\) two-thirds of \(\ge 8\) of the last 12 of 24 weeks, no rescue
ADVANCE IV, efgartigimod Platelet \(\ge 50 \times 10^9\)/L for at least 4 of the last 6 weeks
VAYHIT2, ianalumab Time to treatment failure
Orelabrutinib phase 2 Two consecutive platelet counts \(\ge 50 \times 10^9\)/L within 4 weeks, no rescue
Rilzabrutinib phase 1/2 part A Two consecutive counts \(\ge 50 \times 10^9\)/L and an increase \(\ge 20 \times 10^9\)/L over baseline, no rescue

Two consequences for this project. The threshold the model has to simulate is 50, not the IWG’s 30 or 100, and the durability window differs by trial, so Section 15’s simulated endpoint is a parameter of the simulation rather than a constant. And VAYHIT2’s time-to-treatment-failure endpoint is a different mathematical object from a durable response proportion, which is why Section 16’s first external check compares a predicted plateau rather than a predicted rate.

3.5 Time to response by agent, and where the transduction prior comes from

Table 3 of the IWG paper reports, per agent, the time to initial response and the time to peak response over the reported dose range. It is the published prior on the transduction delay \(k_{e0}\) in Section 8.4, and unlike everything else feeding that prior it has been read.

IWG 2009 Table 3, selected rows ✅
Agent Reported dose range Time to initial response Time to peak response
IVIg 0.4-1 g/kg per dose, 1-5 doses 1-3 d 2-7 d
Anti-D 75 µg/kg iv 1-3 d 3-7 d
Prednisone 1-4 mg/kg po daily, 1-4 wk 4-14 d 7-28 d
Romiplostim (AMG531) 3-10 µg/kg weekly sc 5-14 d 14-60 d
Eltrombopag 25-75 mg po daily 7-28 d 14-90 d
Rituximab 375 mg/m² per dose, 4 weekly doses 7-56 d 14-180 d

The rituximab row is the one this project needs, and its width is the point. Time to peak response spans 14 to 180 days, which is a factor of thirteen. A transduction delay with that much spread across patients is why Section 10 gives up on estimating \(k_{e0}\) and \(\eta_i\) jointly, and why requirement 4 in Section 9 asks for follow-up past the point where a slow response would have completed. The contrast with the two rows above it is also the argument of Section 19: a thrombopoietin receptor agonist reaches peak effect in weeks through a mechanism that acts directly on platelet production, and that is the class whose PK-platelet models are published.

4. Biomarkers

What can be measured in blood on this study, what each one reports on, and what the model does with it. The tissue compartment of uncertainty 1 has no direct measurement; the last three rows are its proxies. Everything in this section outside pharmacometrics is from general knowledge and carries ❌ in the references (entry 30) until sourced.

Blood biomarkers in ITP and what the model does with each
Biomarker What it reports Role in the model Section
Platelet count The endpoint. Destruction and production together The response variable 3, 12
Immature platelet fraction (IPF) Share of circulating platelets with residual RNA; high when the marrow is compensating for destruction Separates a fall in \(k_{\rm dest}\) from a rise in production; one extra number on a routine count 9, 8.3
Serum thrombopoietin (TPO) Inversely related to megakaryocyte and platelet mass; normal or low in ITP, high in marrow failure Diagnostic context; a covariate on baseline production if it varies
Glycocalicin Soluble fragment of platelet GPIb shed on destruction A direct destruction-rate marker; rarely measured in trials
CD19+ or CD20+ B cells Circulating B-cell count; below the flow limit at every therapeutic dose The layer that carries dose into the model, through time to repopulation 6.2, 8.2
B-cell subsets: naive, memory (CD27+), plasmablasts (CD27hi CD38hi) Which compartment is repopulating, and whether plasmablasts, the immediate precursors of antibody-secreting cells, are still being generated The circulating plasmablast count is the closest blood proxy for the tissue compartment 2, 8.5
Serum total IgG, IgM, IgA Bulk antibody, maintained mostly by long-lived plasma cells the drug does not deplete Covariate identifying the plasma-cell floor \(\phi_i\); the whole mechanism for an FcRn blocker 8.5
BAFF B-cell survival factor; rises when B cells are depleted and falls as they return A feedback marker of depletion depth in tissue as well as blood; the target of ianalumab 8.5
Anti-platelet IgG by MAIPA, against GPIIb/IIIa, GPIb/IX, GPIa/IIa The autoantibody itself, by glycoprotein target. Direct assay on the patient’s platelets is more sensitive than indirect on serum. Positive in roughly half to 60% of ITP \(A(t)\) observed rather than latent, where positive and serial 8.3, 8.5
Anti-GPIb/IX specificity Antibodies to GPIb/IX clear platelets by desialylation and hepatic uptake, Fc-independently, and predict poorer response to IVIg and corticosteroids A response covariate, and a mechanism that FcγR-blocking drugs would miss 12
Platelet-associated IgG (PAIgG) Total IgG on platelets; elevated in thrombocytopenia of any cause Not useful; listed so that it is not mistaken for the row above

Autoantibodies that can be measured in circulation. The answer to whether a known autoreactive antibody is measurable is yes, with the qualifications in the MAIPA row: the antibodies are against platelet glycoproteins, GPIIb/IIIa most often, then GPIb/IX, and a minority against GPIa/IIa or GPV; the assay names the target; and it is negative in a large minority of patients with a clinical diagnosis of ITP, either because the titre is low or because the antibody is against something the panel does not carry. A serial MAIPA in the patients who are positive at baseline is the one addition to a standard protocol that would let the model see the autoantibody fall before the platelet count rises (Section 8.5).

5. The drugs this document refers to

Every drug named anywhere in this folder, by mechanism, with what it does in ITP and where it stands. The compound the specification is written for is generic; these are the sources of its parameters and its external checks. First-line therapy, corticosteroids, intravenous immunoglobulin (IVIg) and anti-D, is context only. The last column says whether the status and the response figure were checked against a source; most were not.

Drugs named in this project, by mechanism
Drug Target and mechanism Status in ITP, and line Response as reported
Deplete B cells
Rituximab Chimeric anti-CD20 IgG1; complement- and antibody-dependent lysis of CD20+ B cells Off-label second line, 375 mg/m² or 100 mg weekly × 4 About 60% initial, 20 to 30% at five years
Ianalumab Afucosylated anti-BAFF-R IgG1; ADCC-enhanced depletion plus BAFF-R signalling blockade Phase 3 complete (VAYHIT2, with eltrombopag after first line; VAYHIT1 later line); not approved at time of writing; 4 monthly doses 12-month freedom from treatment failure 54% at 9 mg/kg, 51% at 3 mg/kg, 30% placebo
Obinutuzumab, obexelimab, telitacicept, povetacicept Anti-CD20; anti-CD19/FcγRIIb; TACI-Fc; APRIL/BAFF inhibitor Searched, nothing in ITP found None
Deplete plasma cells
Mezagitamab Anti-CD38 IgG1, subcutaneous; depletes plasmablasts and plasma cells Phase 2b positive, Phase 3 ongoing, chronic ITP 10 of 11 responders at 600 mg weekly × 8, 3 of 13 on placebo
Remove the autoantibody
Efgartigimod FcRn-blocking IgG1 Fc fragment; accelerates IgG catabolism Approved in Japan 2024, IV, primary ITP; ADVANCE NEXT running for FDA Sustained response 25.6% against 6.7% placebo, ADVANCE IV
Rozanolixizumab Anti-FcRn IgG4, subcutaneous Phase 2 in ITP; ITP programme not continued Platelets \(\ge 50\) at least once in 36 to 67% by arm
IVIg Pooled IgG; FcRn competition and FcγR blockade First line and rescue The rodent-model drug of entries 28 and 29
Block the destruction step
Fostamatinib SYK inhibitor prodrug (active moiety R406); blocks FcγR-mediated phagocytosis FDA 2018, EMA 2020; chronic ITP after insufficient response to a prior therapy; 100 mg BID titrated to 150 Stable response 18% across two Phase 3 studies
Rilzabrutinib Reversible covalent BTK inhibitor; less autoantibody production and less FcγR phagocytosis FDA August 2025; persistent or chronic ITP after insufficient response; 400 mg BID Durable response 23% against 0% placebo, LUNA 3
Orelabrutinib BTK inhibitor, once daily Phase 2 in China, Phase 3 ongoing 40% at 50 mg, 22% at 30 mg, 33 patients
Raise platelet production
Romiplostim TPO receptor agonist peptibody, subcutaneous weekly FDA 2008; ITP after insufficient response to corticosteroids, immunoglobulins or splenectomy; titrated 1 to 10 µg/kg Response in most patients while dosed; loses response on stopping
Eltrombopag Oral small-molecule TPO receptor agonist FDA 2008; same line; 25 to 75 mg daily titrated Same pattern; the 155-patient model is entry 26
Avatrombopag Oral TPO receptor agonist FDA 2019 for chronic ITP Same pattern
Lusutrombopag Oral TPO receptor agonist Not an ITP drug; approved for procedures in liver disease Source of a model only

The four mechanism groups are the four rows of the chain in the In brief diagram, read from the drug end. Depleting B cells or plasma cells acts on autoantibody production; removing the autoantibody acts on its concentration; blocking the destruction step acts on what the autoantibody does; and raising production bypasses the chain. The published PK-platelet models are all in the last group, and Section 19 is about why.

6. The question in four parts

“Can the model select the dose” is not answerable as written. It becomes answerable when split into four, each of which the simulation in Section 15 reports separately.

6.1 Is dose to exposure identifiable? Yes, and it is not the problem. A two-compartment model with linear clearance fits from the escalation PK alone.

6.2 Is exposure to B-cell depletion identifiable? The slope is; the \(EC_{50}\) is not, at this study’s size, and the published B-cell depletion models say why. Every one found uses an indirect-response structure with the drug raising the B-cell loss rate, and three report what happened to \(EC_{50}\) (references, entries 31 to 33):

  • ofatumumab, 1,486 patients over doses from 3 mg to 700 mg, two low-dose arms giving partial depletion: \(EC_{50}\) estimated at 8.5% RSE;
  • rituximab, 63 paediatric patients at full dose: \(E_{\max}\) and \(ED_{50}\) both estimated, \(ED_{50}\) at 61% RSE with a 13-fold confidence interval, and the authors attribute the imprecision to “the highly efficacious doses”;
  • inebilizumab, 213 subjects over a 100-fold dose range: the \(E_{\max}\) model was abandoned because the estimated \(EC_{50}\) fell below the PK assay’s limit of quantification, and a log-linear slope replaced it.

The arithmetic behind all three. With loss rate \(k_{\rm out}(1 + E)\) and \(E = E_{\max} C / (EC_{50} + C)\), the nadir is \(B_0 / (1 + E)\). Both \(E_{\max}\) estimates are about 155 to 160, so at \(C = EC_{50}\) the nadir is already about 1% of baseline, below the flow-cytometry limit. Half-maximal depletion happens near \(C = EC_{50} / E_{\max}\), two orders of magnitude below \(EC_{50}\). The whole band of concentrations over which the cell count is informative therefore sits between \(EC_{50}/E_{\max}\) and \(EC_{50}\), far below therapeutic exposure and often below the PK assay. Above that band the depth is censored; repopulation begins when the concentration falls back through it, and its timing identifies \(E_{\max}/EC_{50}\) as one slope. \(E_{\max}\) and \(EC_{50}\) separate only if some cohort sits inside the band with a quantifiable nadir and quantifiable PK. Ofatumumab had both, from its 3 mg and 10 mg arms; inebilizumab, with a wider dose range, had neither.

So the proportional model is the identifiable one, and not by convention. The specification parameterizes Section 8.2 as a log-linear slope on the loss rate and keeps \(E_{\max}\) as an alternative to test only if the lowest escalation cohort shows a quantifiable nadir. The full argument, with the simulations and the design that would change the answer, is the B-cell indirect-response estimation project. The slope is what the time to repopulation depends on, so nothing downstream loses anything.

What the duration signal carries. With linear PK in the tail, each doubling of dose adds one PK half-life to the time the concentration spends above the transition band: about 10 days for ianalumab, 15 to 20 for rituximab. Between-patient variability in the recovery rate was 126% CV in the inebilizumab model and 22% CV in the paediatric rituximab CD19 half-life. At three to six patients per level the dose signal in \(T_{\rm rep}\) is one half-life per doubling against that spread, which is why Section 9 asks for the escalation to span at least eight-fold and for B cells monthly through repopulation, and why Section 10 profiles the slope rather than assuming it. The duration is an input to the model and not a selection criterion; see the last paragraph of the In brief.

6.3 Is B-cell depletion to platelet response identifiable? This is the question. The link is indirect, delayed by weeks, and mediated by two unmeasured states. Section 10 sets out what has to be assumed for it to be identifiable at all. The FDA review of rilzabrutinib (references, entry 22) is an existence proof one step short: a longitudinal platelet model for a drug that blocks destruction fit adequately at 305 patients pooled across two trials at one dose. Whether it does at 40 patients across four doses is what Section 15 tests.

6.4 Does a wrong answer cost anything? The selection rule in Section 13 is scored against a loss function, not against a point estimate. Selecting a dose one level above the true plateau bottom costs exposure and drug supply; selecting one level below costs the Phase 2 result. Those are not symmetric and the rule reflects that.

7. Why depth of B-cell depletion carries no dose information

Peripheral B-cell counts fall below the limit of quantification within days at every dose that any B-cell depleting agent is developed at. The evidence in the references, all unverified:

  • rituximab at 100 mg weekly for four weeks produces a pooled overall platelet response rate close to the rate at 375 mg/m² weekly for four weeks, a roughly four-fold difference in dose, and complete peripheral depletion at both;
  • ianalumab produces “rapid and profound” peripheral depletion across the doses studied in primary Sjögren’s syndrome, which is why its Phase 2b dose selection used tissue receptor occupancy rather than circulating B cells as the second biomarker;
  • in ITP specifically, ianalumab at 3 mg/kg and 9 mg/kg gave 12-month freedom-from-treatment-failure probabilities within a few percentage points of each other in the Phase 3 VAYHIT2 trial, over a three-fold dose range.

Three consequences follow for the model.

Duration replaces depth. The dose-dependent quantity is the time from the last administration until circulating B cells return to a defined fraction of baseline. Call it \(T_{\rm rep}\). It is a derived quantity of the model in Section 8.2, it is directly observable in the escalation data, and it is monotone in dose across a range where depth is flat. It is also a weak signal per patient, one PK half-life per doubling of dose against a between-patient spread in recovery rate that published models put at 22 to 126% CV (Section 6.2), so it is read across the whole escalation and not from any one cohort.

The plateau boundary is the target of estimation. The selection question is where the dose-response curve stops rising. A model built to estimate a slope will be asked for a number the data cannot give; a model built to locate a plateau boundary is being asked something the escalation cohorts can answer.

Tissue depletion carries the dose-response the platelet endpoint depends on, and it is not measured. Autoreactive plasmablasts in ITP disseminate to spleen and bone marrow. Peripheral clearance says nothing about tissue clearance, and no Phase 1 study biopsies spleen. This is the largest source of structural uncertainty in the specification, and Section 17 treats it as the main route to a negative answer.

8. The model

Four layers, written top to bottom in the order the drug traverses them. Each carries a note on whether its parameters are estimated here, fixed from the literature, or collapsed away.

8.0 State variables and parameters of the pharmacodynamic model

The PK layer is standard and not tabulated. Everything downstream of it is, so that the equations in 8.2 to 8.4 can be read against one list. “Full” means Sections 8.2 and 8.3; “reduced” means Section 8.4, the model that gets fit.

Pharmacodynamic states and parameters
Symbol What it is Units Full or reduced How it is set
States
\(B(t)\) Circulating B-cell count cells/µL Both Observed
\(B_{\rm tissue}(t)\) B cells in spleen and marrow relative to baseline Full Unobserved; proxies in Section 4
\(A(t)\) Anti-platelet IgG relative to \(A_{50}\) Full Latent, or observed where MAIPA is serial
\(P_1, P_2, P_3\) Platelet precursors in the maturation chain cells/µL-equivalent Both Unobserved
\(P(t)\) Circulating platelet count \(10^9\)/L Both Observed; the endpoint
\(u(t) = 1 - B/B_0\) Depletion signal none, 0 to 1 Reduced Derived from \(B\)
\(Z(t)\) Transduced depletion signal, \(u\) after a lag none, 0 to 1 Reduced Derived
B-cell layer
\(B_0\) Baseline B-cell count, \(= k_{\rm in}/k_{\rm out}\) cells/µL Both Estimated per patient
\(k_{\rm in}\) B-cell production rate cells/µL/day Both Estimated
\(k_{\rm out}\) B-cell loss rate; sets repopulation speed 1/day Both Estimated; published 0.0124/day (ofatumumab), half-life 173 d (rituximab)
\(S\) Drug effect slope on \(k_{\rm out}\) per log concentration none Both Estimated
\(C_{\rm ref}\) Concentration scale for \(S\) mg/L Both Fixed at PK LLOQ
\(E_{\max}, EC_{50}\) Alternative saturable drug effect none, mg/L Both Fit only if a cohort has a quantifiable nadir (Section 6.2)
\(\alpha\) Repopulation threshold defining \(T_{\rm rep}\) none Both Fixed, 0.2
Autoantibody layer
\(k_{\rm syn}\) Autoantibody production rate at baseline 1/day Full Fixed; scale absorbed into \(A_{50}\)
\(k_{\rm deg}\) IgG elimination rate 1/day Full Fixed, half-life about 21 d
\(\phi_i\) Fraction of autoantibody production from long-lived plasma cells the drug does not deplete none, 0 to 1 Full Random effect; the non-responder fraction
\(A_{50}\) Autoantibody level doubling the destruction rate same as \(A\) Full Fixed; scale
Platelet layer
\(k_{\rm prod}\) Precursor production rate cells/day Both Fixed from TPO-RA models
\(k_{\rm tr}\) Maturation transit rate 1/day Both Fixed from TPO-RA models
\(k_{\rm dest,0}\) Physiological platelet loss rate 1/day Both Fixed from 7 to 10 d lifespan
\(k_{\rm dest}(t)\) Actual platelet loss rate under disease and drug 1/day Both Derived: from \(A(t)\) in 8.3, from \(Z(t)\) in 8.4
Reduced-model link
\(k_{e0}\) Transduction rate from \(u\) to \(Z\) 1/day Reduced Informative prior from rituximab time to response
\(\theta_i\) Untreated excess destruction, as a multiple of \(k_{\rm dest,0}\) none Reduced Estimated per patient from pre-treatment counts
\(\eta_i\) Fraction of \(\theta_i\) the drug can remove; \(\approx 1 - \phi_i\) none, 0 to 1 Reduced Random effect with a mixture; the responder indicator

8.1 Pharmacokinetics

Two compartments with linear clearance. Writing \(C\) for concentration in the central compartment and \(C_p\) in the peripheral,

\[ V_c \frac{dC}{dt} = -CL \cdot C - Q(C - C_p) + \text{input}, \qquad V_p \frac{dC_p}{dt} = Q(C - C_p) \]

No target-mediated term. B-cell-mediated elimination is real for anti-CD20 antibodies and the published models carry it (references, entries 32 and 34), but at this study’s size it is not estimable and it is not the question. Where a patient’s baseline B-cell count shifts clearance, it enters as a covariate on \(CL\), which is also how Section 11 handles the confounding it causes. Estimated here, with informative priors on \(CL\) and \(V_c\) from typical immunoglobulin G1 values.

8.2 Circulating B cells

An indirect response model on the loss rate, with the drug effect written as a log-linear slope, which is the form the published models converge on when depletion is complete at every studied dose (Section 6.2):

\[ \frac{dB}{dt} = k_{\rm in} - k_{\rm out}\left(1 + S \cdot \log(1 + C/C_{\rm ref})\right) B \]

with \(C_{\rm ref}\) fixed at the PK assay’s limit of quantification so that \(S\) is in units the data can see. The \(E_{\max}\) form, \(E = E_{\max} C/(EC_{50} + C)\), is the alternative to fit only if the lowest escalation cohort has a nadir above the flow-cytometry limit; when it does not, \(E_{\max}\) and \(EC_{50}\) are identified only as their ratio and the slope form is the same model with one fewer parameter.

with \(B(0) = B_0 = k_{\rm in}/k_{\rm out}\). The quantity carried forward is

\[ T_{\rm rep} = \inf\{\,t > t_{\rm last} : B(t) \ge \alpha B_0\,\}, \qquad \alpha = 0.2 \]

the time from the last dose until circulating B cells recover to 20% of baseline. \(\alpha\) is a property of the definition, not of a dataset, and it is fixed at 0.2 so that \(T_{\rm rep}\) is measurable above the flow cytometry quantification limit in most patients.

Estimated here with the caveat of Section 6.2: \(S\) is estimated, and \(k_{\rm out}\) — which sets the repopulation rate and therefore \(T_{\rm rep}\) — is the parameter the escalation data informs best. Published values to compare against: \(k_{\rm out}\) 0.0124 per day for ofatumumab, a CD19 half-life of 173 days for paediatric rituximab, a B-cell lifespan of 391 days for inebilizumab (references, entries 31 to 33).

8.3 Autoantibody and platelet dynamics

The published platelet-side structure comes from the thrombopoietin receptor agonist (TPO-RA) literature, where a four-compartment lifespan model is standard: one precursor production compartment, two transit compartments for maturation, and one circulating platelet compartment. Written with \(P_1, P_2, P_3\) for the maturation chain and \(P\) for circulating platelets,

\[ \frac{dP_1}{dt} = k_{\rm prod} - k_{\rm tr} P_1, \qquad \frac{dP_i}{dt} = k_{\rm tr}(P_{i-1} - P_i), \qquad \frac{dP}{dt} = k_{\rm tr} P_3 - k_{\rm dest}(t)\, P \]

ITP enters through the destruction rate, which anti-platelet IgG \(A(t)\) elevates above its physiological value. \(A(t)\) is latent unless an anti-platelet antibody assay was run, and Section 8.5 says what changes when it was:

\[ k_{\rm dest}(t) = k_{\rm dest,0}\left(1 + \frac{A(t)}{A_{50}}\right), \qquad \frac{dA}{dt} = k_{\rm syn}\, g(B_{\rm tissue}) - k_{\rm deg} A \]

Fixed from the literature, not estimated. \(k_{\rm tr}\) and \(k_{\rm dest,0}\) come from the TPO-RA models and from the 7 to 10 day physiological platelet lifespan. The eltrombopag ITP model (references, entry 26) made the same decision for the same reason: its authors could not estimate the production rate and the transit rate from 155 ITP patients and fixed both to healthy-volunteer values, 1.43 Gi/L·h and 0.0253 h⁻¹. \(k_{\rm deg}\) is the IgG elimination rate, roughly a three-week half-life. Nothing in an ITP dataset without a TPO-RA arm perturbs platelet production enough to identify the maturation chain independently.

The reason ITP responses to B-cell depletion take four to eight weeks, and the reason a substantial minority never respond at all, both live in the \(A\) equation. Long-lived plasma cells carry neither CD20 nor the B-cell activating factor receptor, are not depleted by these agents, and continue to secrete \(A\). Represent that as a floor:

\[ k_{\rm syn} g(B_{\rm tissue}) = k_{\rm syn}\left[(1 - \phi)\, \frac{B_{\rm tissue}}{B_{\rm tissue,0}} + \phi \right] \]

with \(\phi_i\) the fraction of autoantibody production in patient \(i\) that survives complete B-cell depletion. \(\phi_i\) is what separates a responder from a non-responder, it is a between-patient random effect, and it is not identified from an eight-week observation window. Its image in the reduced model of Section 8.4 is \(\eta_i \approx 1 - \phi_i\), and Section 12 makes that the mixture parameter.

8.4 The reduced model a small dataset can support

Sections 8.2 and 8.3 chained together carry more parameters than fifty patients will support. The model that gets fit collapses the tissue and autoantibody layers into one delayed signal, defined here in words before the equation.

The depletion signal \(u(t)\). \(u(t) = 1 - B(t)/B_0\), where \(B\) is the circulating B-cell count and \(B_0\) its baseline. It is 0 before dosing, 1 when B cells are fully depleted, and returns toward 0 as they repopulate. It is what the B-cell layer hands to the platelet layer, and it is observed.

The transduced signal \(Z(t)\). \(u(t)\) passed through a first-order lag with rate \(k_{e0}\):

\[ \frac{dZ}{dt} = k_{e0}\,\big(u(t) - Z\big), \qquad Z(0) = 0 \]

\(Z\) lies between 0 and 1 and follows \(u\) with a delay of about \(1/k_{e0}\). It stands in for everything between depletion of circulating B cells and a change in platelet destruction that the study does not measure: depletion in spleen and marrow, the turnover of plasmablasts, and the decay of the autoantibody they made. A \(k_{e0}\) half-life of about six weeks reproduces rituximab’s time to response (Section 3.5). \(Z\) is not the drug effect on B cells, which Section 8.2 writes as \(S \cdot \log(1 + C/C_{\rm ref})\); it is one layer further down.

How \(Z\) moves the destruction rate. \(k_{\rm dest}(t)\) is the loss rate of circulating platelets in the last equation of Section 8.3, \(dP/dt = k_{\rm tr} P_3 - k_{\rm dest}(t)\, P\). The platelet chain is unchanged from 8.3; the reduced model changes only what drives \(k_{\rm dest}\), replacing the latent autoantibody \(A(t)\) with \(Z(t)\):

\[ k_{\rm dest}(t) = k_{\rm dest,0}\left(1 + \theta_i \left[1 - \eta_i Z(t)\right]\right) \]

Before dosing, \(Z = 0\) and \(k_{\rm dest} = k_{\rm dest,0}(1 + \theta_i)\): the patient’s untreated excess destruction, with \(\theta_i\) set by the pre-treatment platelet count. At full transduced depletion, \(Z = 1\) and \(k_{\rm dest} = k_{\rm dest,0}(1 + \theta_i(1 - \eta_i))\): the drug has removed a fraction \(\eta_i\) of the excess. A good fit has \(\theta_i\) of several-fold, since untreated ITP platelet counts are a small fraction of normal, and \(\eta_i\) near 1 in a responder and near 0 in a non-responder.

Three parameters remain on the drug side: \(k_{e0}\), taking an informative prior from published time to platelet response after rituximab in ITP; \(\theta_i\), identified from the pre-treatment platelet count; and \(\eta_i \in [0, 1]\), the fraction of that severity the drug can remove in patient \(i\). \(\eta_i\) is the mixture. Its image in the full model is \(1 - \phi_i\), where \(\phi_i\) (Section 8.3) is the fraction of the patient’s anti-platelet antibody that is made by long-lived plasma cells, which carry neither CD20 nor BAFF-R and are not depleted; a patient with \(\phi_i\) near 1 keeps making the antibody after every B cell is gone, and is the non-responder.

This mixture has a published name and a precedent. The eltrombopag ITP model (references, entry 26) put a mixture on the drug-effect slope, estimated 81% responders with 19% at slope zero, and then used the fitted model to simulate response-guided titration. Section 8.4 is that model with the drug effect moved from production to destruction and the lag lengthened from two weeks to six.

This reduction is the specification’s main structural assumption, and Section 15 tests it against data generated from the unreduced model. If the reduced model recovers the right dose from data generated by the full one, the reduction is safe at this sample size. If it does not, the answer to the project’s question is no for the reason given in Section 17.

9. What the dataset has to contain

The model above is a set of requirements on the study, and a study that does not meet them cannot be rescued by analysis. Stated as requirements so that they can be checked against a real protocol before any fitting.

Data requirements, in decreasing order of consequence
# Requirement Why the model needs it If absent
1 At least three dose levels with \(\ge 3\) patients each Any dose-response at all Fatal
2 A top dose at least 4-fold above the bottom dose \(T_{\rm rep}\) separation between levels Fatal
3 B-cell counts weekly for 8 weeks, then monthly to repopulation \(k_{\rm out}\), and therefore \(T_{\rm rep}\) Fatal
4 Follow-up past B-cell repopulation in most patients \(T_{\rm rep}\) is right-censored otherwise Severe
5 Platelet counts at every visit, 6 months minimum Separating responders from non-responders Fatal
6 Pre-treatment platelet counts, \(\ge 2\) per patient \(\theta_i\), the baseline severity Severe
7 Rescue therapy dates and doses Platelet rises not attributable to drug Severe
8 Concomitant TPO-RA and corticosteroid exposure recorded The largest competing platelet driver Severe
9 Serum total IgG at every visit Identifies the plasma-cell floor \(\phi_i\) (Section 8.5) Valuable
9a Anti-platelet IgG by MAIPA, serially, where positive at baseline Makes \(A(t)\) observed rather than latent (Section 8.5) Valuable
10 Immature platelet fraction Separates production from destruction Valuable

Requirement 10 deserves the argument it gets in Section 12. Immature platelet fraction (IPF) is the proportion of circulating platelets with residual RNA, measurable on a standard haematology analyser, and it stands to thrombopoiesis as the reticulocyte count stands to erythropoiesis. In ITP it is characteristically elevated, because destruction is high and the marrow is compensating. A drug that lowers destruction should lower IPF while raising platelet count, and a drug that raised production would raise both. That is a structural identifiability question answered by one extra number on a blood count that was drawn anyway.

10. Identifiability

Three parameters in Section 8.4 compete to explain the same feature of the data, which is that some patients’ platelet counts rise weeks after dosing and others’ do not.

\(k_{e0}\) against \(\eta_i\). A slow transduction rate and a partial drug effect both flatten the platelet trajectory. They separate only if follow-up extends past the point where a slow response would have completed, which is requirement 4 in Section 9 restated on the platelet side. Fix \(k_{e0}\) from the literature prior and estimate \(\eta_i\); the reverse is not supportable.

\(\eta_i\) against \(\theta_i\). Baseline severity and drug effect are separated by the pre-treatment platelet counts and by nothing else, which is requirement 6.

\(EC_{50}\) against \(E_{\max}\) in the B-cell layer. Not separable at doses that all saturate, for the reason in Section 6.2, and the model is parameterized as a slope so that the question does not arise. It arises again only if the lowest escalation cohort has a quantifiable nadir, which is an argument for a low starting dose that a pure safety argument would not make. The diagnostic that replaces the \(EC_{50}\) profile is a profile likelihood on the slope \(S\), and it is expected to be wide.

Report all three as identifiability diagnostics on every simulated fit, not as prose: condition number of the Fisher information matrix, the profile likelihood for the slope \(S\) and for \(k_{\rm out}\), and the correlation between the \(\eta_i\) and \(\theta_i\) empirical Bayes estimates.

11. Exposure-response confounding

An exposure-response analysis of an escalation-plus-expansion study is a known generator of false positives, and the review literature on time-dependent clearance in therapeutic antibodies names the two mechanisms that apply here.

Target burden drives clearance. B cells are the target sink. A patient with a high B-cell burden clears the drug faster and therefore has both lower exposure and, plausibly, more disease to remove. Exposure and outcome then correlate for a reason that has nothing to do with dose. The PK model in Section 8.1 is linear and does not carry this mechanism; baseline B-cell count is tested as a covariate on clearance so that the correlation is named rather than left in the residual, and the selection rule below is written on dose rather than exposure so that it cannot act on the decision.

An expansion with few arms is close to the worst case. The published review finding is that designs with one tested dose level have a high probability of a false-positive exposure-response result, because all the exposure variation is between-patient and between-patient exposure variation is confounded with everything else about the patient. Two to four arms are better than one and still leave most of the exposure spread inside arms.

The consequence for the analysis plan is a rule, not a caution. Dose, not exposure, is the independent variable in the dose selection rule of Section 13. Individual exposure enters the model as the thing dose produces, and the selection is made by simulating dose levels forward from the fitted model rather than by regressing outcome on observed area under the concentration curve. An exposure-response plot may be produced as a diagnostic; it does not select a dose.

12. Choice of endpoint for fitting

The Phase 2 and Phase 3 endpoints in ITP are dichotomies. Durable response as platelet count \(\ge 50 \times 10^9\)/L over a specified share of weeks is what gets reported and what a label carries. Fitting that dichotomy directly at Phase 1 sample sizes wastes most of the information in the dataset, and the statistical literature on dichotomizing continuous outcomes quantifies the loss.

The rilzabrutinib FDA review (references, entry 22) shows the two endpoints disagreeing on the same 305 patients: the continuous change in platelet count rose with exposure, and the durable-response dichotomy was flat across exposure quartiles. The fostamatinib reviews show the same thing at 79 patients, where a slope at Week 12 had vanished by the Week 14 to 24 window.

Fit the continuous longitudinal platelet count; derive the dichotomy from the fitted model by simulation. The mixture in Section 8.4 does the work that a responder analysis does by hand: \(\eta_i\) near 0 is a non-responder, and the population distribution of \(\eta\) is what a response rate is a one-number summary of. Simulating the Phase 2 endpoint from the fitted model then gives a predicted response rate per dose with an interval, which is what a Phase 2 decision actually needs.

Two model features this requires and a plain responder analysis does not:

  • A distribution for \(\eta\) with mass near zero. A normal random effect on a logit scale will not reproduce a bimodal response. Candidates are a two-component mixture and a beta distribution with both shape parameters below one. Fit both; report both.
  • Rescue therapy as informative censoring. A patient rescued for a low platelet count is not missing at random. Model the rescue as an event with a hazard driven by the current platelet count, or exclude post-rescue observations and accept the bias, and say which.

13. The dose selection rule

Stated in advance, so that the simulation in Section 15 scores something fixed.

Select the lowest dose whose posterior predictive Phase 2 response rate is within 5 percentage points of the maximum over the dose range studied, with at least 80% posterior probability.

Three properties of that rule, each deliberate.

It targets the bottom of the plateau rather than the maximum, which is the quantity Section 7 argues the data can locate. It uses the predicted clinical endpoint rather than \(T_{\rm rep}\), so that the biomarker’s role stays limited to carrying dose information into the model. And it is asymmetric, because 5 percentage points below the maximum is a tolerance on the downside only.

The comparator rules the simulation scores this against:

Dose selection rules compared in Section 15
Rule Basis
A. Highest safe dose The default when no model is built
B. Lowest dose with \(T_{\rm rep}\) within 10% of maximum Biomarker-only, no platelet model
C. The rule above The model this specification proposes
D. Empirical best observed response rate What the raw data says, no model

Rule A is the baseline. If C does not beat A by enough to change a decision, the project’s answer is that the model should not be built, and that is a publishable finding.

14. How many patients, with and without the model

The question in Section 6 has a cost attached. If the model cannot be built, the dose is chosen from response rates per arm, and this section says what that takes. If the model can be built, it earns its place by needing fewer patients than that, and this section bounds how many fewer. The numbers are computed here rather than quoted, and the Sizing Studies page carries the general method they follow.

14.1 Without a model: picking the best arm from response rates

Two framings, and they give answers an order of magnitude apart.

Detecting a difference between two arms is the hypothesis-test framing, and it is the wrong question for dose selection, but it is the number people reach for first. Two-sided \(\alpha = 0.05\), 80% power:

pairs <- data.frame(p_low = c(.30, .40, .50, .50), p_high = c(.50, .60, .60, .55))
pairs$n_per_arm <- sapply(seq_len(nrow(pairs)), function(i)
  ceiling(power.prop.test(p1 = pairs$p_low[i], p2 = pairs$p_high[i], power = .8)$n))
knitr::kable(pairs, col.names = c("Lower arm", "Higher arm", "Patients per arm"),
             caption = "Patients per arm to detect a difference in durable response rate")
Patients per arm to detect a difference in durable response rate
Lower arm Higher arm Patients per arm
0.3 0.50 93
0.4 0.60 97
0.5 0.60 388
0.5 0.55 1565

Selecting the best of \(K\) arms is the right framing. The design does not need to prove the top arm is better; it needs to pick it with acceptable probability. Probability of correct selection is the chance that the arm with the highest true rate also has the highest observed count, ties split at random, when it beats every other arm by \(\delta\):

set.seed(20260909)
pcs <- function(K, n, p_best, delta, nsim = 20000) {
  p <- c(p_best, rep(p_best - delta, K - 1))
  x <- matrix(rbinom(K * nsim, n, p), nrow = nsim, byrow = TRUE)
  mx <- apply(x, 1, max)
  mean((x[, 1] == mx) / rowSums(x == mx))
}
grid <- expand.grid(K = c(2, 3, 4), delta = c(.10, .15, .20), n = c(10, 20, 30, 40, 60, 80))
grid$pcs <- mapply(pcs, grid$K, grid$n, MoreArgs = list(p_best = .55), delta = grid$delta)
need <- aggregate(n ~ K + delta, data = grid[grid$pcs >= .80, ], FUN = min)
wide <- reshape(need, idvar = "delta", timevar = "K", direction = "wide")
names(wide) <- c("Margin of best arm", "2 arms", "3 arms", "4 arms")
wide[[1]] <- sprintf("%.0f points", 100 * wide[[1]])
knitr::kable(wide, caption = "Patients per arm for 80% probability of picking the best arm (best arm 55%)")
Patients per arm for 80% probability of picking the best arm (best arm 55%)
Margin of best arm 2 arms 3 arms 4 arms
1 10 points 40 80 NA
3 15 points 20 40 60
6 20 points 10 20 30

A blank cell means no value up to 80 per arm reached 80%.

Read against the study in Section 1, which has two expansion arms of 20 to 40. The two-arm column is the one that applies: 15 points apart is picked correctly four times in five at 20 per arm, and 10 points apart needs 40 per arm for the same. Ten points is what the ianalumab 3 against 9 mg/kg comparison and the rituximab low against standard dose comparison both look like (Section 7), so the study as designed sits at the edge of what its raw response rates can decide. The three- and four-arm columns say what adding arms would cost, and it is a lot.

Locating a plateau by comparing arms is not possible at any Phase 1/2 size. The plateau’s defining property is that adjacent arms differ by little. Two arms 5 points apart need

ceiling(power.prop.test(p1 = .50, p2 = .55, power = .8)$n)
[1] 1565

patients per arm to separate. The empirical strategy therefore cannot find the bottom of the plateau; it can only pick the highest tolerated dose and call that the answer, which is comparator rule A in Section 13.

14.2 With a model: where the patients are saved, and where they are not

A model changes the arithmetic in two places and leaves it alone in a third.

Pooling the escalation cohorts and both expansion arms through one dose-response shape. Comparing the two expansion arms uses those two arms’ patients. Fitting a monotone dose-response through the escalation cohorts as well uses every patient in the study to estimate the contrast between the two expansion doses. For five dose levels at two-fold spacing, 4 patients per level from the escalation, and the expansion sitting on the top two levels, the variance of the estimated difference between the two expansion doses under a logistic dose-response linear in log dose, by Fisher information and the delta method, compared to the same difference from the two expansion arms alone:

rel_eff <- function(p_top, delta, x, n) {
  K <- length(x); logit <- function(p) log(p / (1 - p))
  b <- (logit(p_top) - logit(p_top - delta)) / (x[K] - x[K - 1])
  a <- logit(p_top) - b * x[K]
  p <- plogis(a + b * x); w <- n * p * (1 - p)
  info <- rbind(c(sum(w), sum(w * x)), c(sum(w * x), sum(w * x^2)))
  g <- c(p[K] * (1 - p[K]) - p[K - 1] * (1 - p[K - 1]),
         x[K] * p[K] * (1 - p[K]) - x[K - 1] * p[K - 1] * (1 - p[K - 1]))
  v_model <- drop(t(g) %*% solve(info) %*% g)
  v_pair  <- p[K] * (1 - p[K]) / n[K] + p[K - 1] * (1 - p[K - 1]) / n[K - 1]
  v_pair / v_model
}
x <- 0:4                                    # five levels, two-fold spacing
grid <- expand.grid(n_exp = c(20, 30, 40), delta = c(.05, .10, .15, .20))
grid$factor <- mapply(function(ne, d) rel_eff(.55, d, x, c(4, 4, 4, 4 + ne, 4 + ne)),
                      grid$n_exp, grid$delta)
wide <- reshape(grid, idvar = "delta", timevar = "n_exp", direction = "wide")
names(wide) <- c("Margin between expansion doses", "20 per arm", "30 per arm", "40 per arm")
wide[[1]] <- sprintf("%.0f points", 100 * wide[[1]])
knitr::kable(wide, digits = 1,
             caption = "Information gained on the expansion contrast by fitting all cohorts through one logistic shape, as a factor")
Information gained on the expansion contrast by fitting all cohorts through one logistic shape, as a factor
Margin between expansion doses 20 per arm 30 per arm 40 per arm
1 5 points 6.4 5.0 4.2
4 10 points 5.2 4.1 3.5
7 15 points 3.8 3.1 2.6
10 20 points 2.7 2.3 2.0

That factor is the pooling gain, and three things about it. It is large because a two-parameter shape over a sixteen-fold dose range is a strong assumption, and the escalation cohorts pin the shape at the bottom where the expansion has no patients. It falls as the expansion grows, because the expansion arms then carry most of the information anyway. And it holds only if the dose-response has the assumed shape: a plateau shape with a third parameter pools less, and a model with the wrong shape borrows strength from cohorts that are not informative about the contrast and returns a confident wrong answer. Section 15’s wrong-structure case is the check on that, and it is where the factor above gets replaced by a measured one.

Using every platelet count instead of one responder call. The responder dichotomy discards the trajectory. The rilzabrutinib FDA review shows the continuous change in platelet count carrying an exposure signal that the dichotomy did not (references, entry 22), and the augmented-binary literature puts the gain from the continuous component at a third or so of the sample size (entry 20). In this model the gain is bounded, and the bound is the third place.

The Bernoulli ceiling. If a patient is a responder or is not, and the model’s \(\eta_i\) is near 0 or near 1, then once the responder status is known the trajectory adds nothing to the dose-response estimate. The information per patient is that of one coin flip with probability \(P(\eta_i \approx 1 \mid \text{dose})\), and no model changes that. What the continuous data does is classify the coin flip with fewer errors than a threshold on noisy counts does, and shorten the follow-up needed to classify it. Where the mixture is sharp, the model’s whole gain is the pooling factor above; where the mixture is soft, with graded partial responders, the trajectory carries dose information of its own and the gain is larger. Which of those describes B-cell depletion in ITP is not known, and the rituximab response distributions in entry 9 are where to look.

14.3 The scoping table

Putting the pieces together for the study in Section 1, two expansion arms with the better one at 55% and the margin between them at \(\delta\):

What the model buys, by question, in patients per arm
Question Without the model With the model, shape correct With the model, shape wrong
Pick the better arm, \(\delta\) = 20 10 per arm 10 per arm; nothing to gain Same
Pick the better arm, \(\delta\) = 15 20 per arm Under 20 per arm Worse than without
Pick the better arm, \(\delta\) = 10 40 per arm About 20 per arm, from pooling with the escalation Worse than without
Locate the plateau bottom, \(\delta\) = 5 Not possible Only as a shape parameter, and Section 15 says whether it is identified Confidently wrong

The middle column applies the pooling factor to the selection table with a discount for the plateau shape pooling less than the logistic, and the simulation in Section 15 replaces it with a measured number. The right-hand column is the case Section 17 is written for.

The scoping verdict. ⚠️ Two expansion arms of 20 to 40 can pick between doses 15 or more points apart without any model. Between doses 10 points apart, 40 per arm does it on the raw rates and 20 per arm does it only with a model whose dose-response shape is right, so the model is what makes the smaller study sufficient. Locating a plateau bottom is beyond both, except as a parameter of an assumed shape. Build the model when the two candidate doses are expected to be close, and build it only after the simulation in Section 15 has shown that its shape assumption survives being wrong.

15. The simulation study

This is the work. Everything above is the structure it needs.

15.1 Generate. Simulate escalation-plus-expansion datasets from the full model of Sections 8.1 to 8.3, not the reduced model of 4.4. Vary, in a factorial:

  • number of dose levels (3, 4, 5) and patients per level (3, 6);
  • expansion size (20, 40) and which dose it sits at;
  • true shape of the dose-response on \(\eta\): plateau at the second dose level, plateau at the fourth, monotone with no plateau in range;
  • the long-lived plasma cell floor \(\phi\), which sets the ceiling on response rate;
  • follow-up duration (6, 12 months).

15.2 Fit. The reduced model of Section 8.4, with the priors named there, to every simulated dataset. Record convergence, the identifiability diagnostics of Section 10, and the posterior predictive response rate per dose.

15.3 Score. For each of the four rules in Section 13:

  • probability the selected dose equals the true plateau bottom;
  • probability the selected dose is below it, which is the costly error;
  • expected number of dose levels above it, which is the cheap error;
  • the same three under the wrong-structure case, where data came from a model with tissue depletion lagging peripheral depletion and the fitted model assumes they coincide.

15.4 Report. The factorial has too many cells for a figure each. The output is a single figure with true plateau position on the horizontal axis and probability of correct selection on the vertical, one line per rule, one panel per sample size. The tables are the supporting detail.

16. External checks

Two published results the model should be able to reproduce, and neither is used in fitting.

VAYHIT2, ianalumab at 3 and 9 mg/kg in ITP. The 12-month freedom-from-treatment-failure probabilities reported for the two doses differ by a few percentage points, against a larger gap to placebo. A model fit to escalation data covering that range should predict a plateau below 3 mg/kg. If the model instead predicts a meaningful gain from 9 mg/kg over 3 mg/kg, it is wrong in the direction that costs patients drug exposure for nothing, and the reason it is wrong is diagnostic.

Low-dose against standard-dose rituximab in ITP. Pooled response rates for 100 mg weekly and 375 mg/m² weekly are close, across a roughly four-fold dose range, in a meta-analysis of a few hundred patients. Same test, different compound and a much longer literature.

Both checks are on the plateau’s existence, not its exact position, which is the strongest claim aggregate published data can support.

17. What would make the answer no

Written before the work, so that a negative result is recognizable rather than explained away.

Tissue depletion does not track peripheral depletion. If the dose-response the platelet endpoint depends on is in the spleen, and the spleen is invisible, peripheral B-cell kinetics is a biomarker of the wrong compartment, and no amount of platelet modelling recovers the missing dose axis. This is the most likely route to a negative answer and Section 15.3’s wrong-structure case is what detects it.

The response mixture is not identifiable at \(n \approx 40\). If the posterior for the \(\eta\) distribution is dominated by its prior across every simulated dataset, the predicted response rates per dose carry no information that the raw response counts did not, and rule C reduces to rule D.

The plateau is outside the studied range. If the true dose-response is still rising at the top escalation dose, every rule selects the top dose and the model adds nothing. The simulation includes this case so that the frequency with which it can be detected is reported, which is a different and more useful quantity than the frequency with which it is handled.

Concomitant therapy dominates. Most ITP patients entering an escalation study are on a TPO-RA or corticosteroids or both. If the platelet trajectory is mostly explained by those, the drug effect is a small perturbation on a large signal and the sample size required moves out of reach.

18. Milestones

  1. Read the platelet-side literature and fix the platelet parameters. Entries 1 to 3 of the reading queue. Ends with a table of \(k_{\rm tr}\), \(k_{\rm dest,0}\) and platelet lifespan with a source per number.
  2. Read the B-cell-side literature and fix the depletion parameters. Entries 4 to 6. Ends with \(k_{\rm out}\), a prior on \(EC_{50}\), and a plausible \(T_{\rm rep}\) against dose curve.
  3. Set the prior on \(k_{e0}\). Entry 7, the rituximab ITP time-to-response literature. One number and an interval.
  4. Build the full generative model and show it reproduces ITP. Untreated platelet counts in the observed range, a response in four to eight weeks, a response rate in the published range at a saturating dose.
  5. Build the reduced model and fit it to one simulated dataset. The identifiability diagnostics of Section 10 on a single case before any factorial.
  6. Run the factorial and produce the Section 15.4 figure.
  7. Run the external checks of Section 16.

Milestones 1 to 3 are reading and produce a parameter table. Milestone 4 is where the specification first meets contact with reality, and Sections 8 and 9 are expected to change there.

19. Why no such model has been published

Platelet count is the primary endpoint of every ITP trial, so the absence of a PK-platelet model for any immunomodulatory agent in ITP needs an explanation before the project assumes it can build one. The reasons split into two groups: those that would bite this project, and those that explain other programmes and do not apply here. The evidence for each is in references.qmd.

19.1 Reasons that apply to this project

1. The response is binary, and dose moves the fraction of patients who respond rather than how far a responder goes. When the drug works in a patient, the platelet count returns to normal. It does not become more normal at a higher dose. What a higher dose can do is convert more patients from non-responder to responder, so the dose-response lives in the mixture proportion and not in the size of any patient’s rise. Continuous \(E_{\max}\) machinery on the platelet count fits the wrong shape, and the information per patient about dose is that of one binary outcome, which at Phase 1 sample sizes is little to fit. This is the reason the specification writes the drug effect as the mixture \(\eta_i\) in Section 8.4 and scores dose selection on the predicted responder fraction in Section 13.

2. The endpoint is a dichotomy of a dichotomy. CR and R are thresholds on platelet count, and recent trials add a durability requirement over 8 to 12 weeks on top. What is reported is a proportion at one dose, and a proportion at one dose is not something a model is needed for. Section 12 is the response: fit the continuous count and derive the dichotomy from it.

3. Where a continuous platelet model was attempted for a drug that blocks destruction, it failed, except once at ten times this study’s size. The record is in the table below. The one success, rilzabrutinib, had 305 patients pooled across two trials at essentially one dose, and its dichotomous endpoint was flat across exposure even as the continuous count rose with it, which is reason 1 seen in data. The two failures, fostamatinib and efgartigimod, had Phase 3 and Phase 2 sample sizes respectively, both larger than the study in Section 1.

Where a PK-platelet model has been tried in ITP, and what happened
Drug, source Data What was attempted Outcome Entry Reference
Blocking destruction, the project’s class
Fostamatinib, FDA 2018 and EMA 2019 79 ITP patients, two Phase 3 studies, 200 or 300 mg/day titrated Semi-mechanistic, direct-response and indirect-response models on continuous platelet count All failed. Logistic regression at Week 12 found a slope that was gone by the Week 14 to 24 window; FDA read it as a relationship, CHMP as none 23, 24 FDA NDA 209299 multidisciplinary review, 2018, pp. 216-217; EMA/CHMP/654949/2019, pp. 53, 76
Efgartigimod, PMDA 2024 Phase 2 Study 1603, 5 and 10 mg/kg Total IgG to platelet count model, for Phase 3 dose selection Could not explain between-patient variation in platelet count; Phase 3 dose chosen on IgG reduction instead 25 PMDA review report, Vyvgart for ITP, 6 March 2024, §6.R.1
Rilzabrutinib, FDA 2025 305 patients, 10,652 weekly counts, LUNA 2 and 3, essentially one dose Indirect exposure-response model on longitudinal platelet count Fit adequately. Continuous count rose with exposure; durable-response dichotomy flat across exposure quartiles; built after the dose was chosen, to confirm it 22 FDA NDA 219685 integrated review, 2025, §14.5.2
Raising production, for contrast
Eltrombopag, Hayes 2011 155 ITP patients, one 6-week Phase 3 Lifespan model with a 19% non-responder mixture Fit; used to simulate response-guided titration 26 Hayes et al., J Clin Pharmacol 2011;51:1403, doi:10.1177/0091270010383019
Romiplostim, ASH 2021 268 ITP patients, 7 studies, 1 to 15 µg/kg Progenitor, four transit, circulating; effect linear in exposure plus cumulative dose Fit; used to confirm titration 27 Blood 2021;138(Suppl 1):4221, ASH
Blocking destruction, in animals
IVIg, Balthasar group 2003 and 2007 Rat and mouse antibody-induced ITP FcRn-competition PK with indirect response from antibody concentration to platelet count Fit; about half of the IVIg effect attributed to faster antibody clearance 28, 29 Hansen and Balthasar, J Pharm Sci 2003;92:1206, PubMed 12761810; Deng and Balthasar, J Pharm Sci 2007;96:1625, doi:10.1002/jps.20828

What the three regulatory attempts say together. At Phase 2 size the model did not identify; at Phase 3 size with two titrated doses no continuous model fit; at pooled Phase 2 and 3 size with one dose a model fit and showed a flat dose-response on the endpoint that counts. None was used to pick a dose. What remains open is whether a model built for the mixture, with the platelet layer fixed and the escalation cohorts pooled through a dose-response shape (Sections 8 and 14), does better at the sample sizes those three had. Section 15 is that test.

19.2 Reasons that explain other programmes and do not apply here

Recorded because they account for most of the literature’s shape, and so that they are not mistaken for obstacles to this project.

4. The mechanism that got modelled has a graded dose-response. A thrombopoietin receptor agonist raises \(k_{\rm prod}\) directly, every patient’s count moves, the size of the move scales with dose, and the loop closes in two weeks. That is a textbook indirect-response problem, and the eltrombopag, romiplostim, avatrombopag and lusutrombopag models exist because of it. This project’s drug class acts on \(k_{\rm dest}\) through an unmeasured autoantibody, with a peak-response lag of 14 to 180 days (Section 3.5). The contrast explains which drugs have models; it is not itself an obstacle, since the project does not expect a graded response.

5. Doses are titrated to the endpoint. Thrombopoietin receptor agonists and fostamatinib are dose-adjusted to hold platelets in a window, so dose depends on response and a standard exposure-response analysis reads backwards. The study in Section 1 gives fixed doses in fixed arms, so this does not arise, and the fostamatinib failure in the table above is partly this reason rather than reason 1.

6. Dose-ranging designs in ITP have destroyed the dose contrast. Rilzabrutinib’s dose-finding used intrapatient escalation over four levels, confounding dose with time on drug in every patient. The design in Section 1 is parallel arms, so this does not arise either. The orelabrutinib Phase 2 randomized two doses and is the one published ITP design of the shape assumed here.

7. Sponsors fell back to the clean biomarker, and regulators accepted it. Efgartigimod’s Phase 3 dose was chosen on IgG reduction after the platelet model failed; ianalumab’s Phase 2b doses in Sjögren’s on B cells and tissue receptor occupancy; rilzabrutinib’s 400 mg BID on BTK occupancy plateaus plus raw Part A response by dose. In every case the fast, continuous, saturable biomarker made the decision. This is the path the project declines to take, for the reason in the In brief: complete depletion of circulating B cells is not the same thing as maximal platelet response.

20. The intended use

A pharmacometrician holding escalation and expansion data on a B-cell depleting agent in ITP, being asked for a Phase 2 dose. This document says what such a model can and cannot deliver, what the study has to have measured for it to be buildable, and how often it beats picking the highest safe dose.

The expected finding, which Section 15 tests rather than assumes, is that the model locates a plateau boundary the raw response counts cannot resolve at \(n \approx 40\), and that it does so only where B-cell repopulation has been followed to completion.

Back to top