Pharmacometrics Modeling Skills
Scoping a small set of AI skills for building and evaluating pharmacometrics models
Idea, 2026-10-02. Nothing is built. The vision, governance and review rules are in the proposal; this page carries the working detail, starting with the first skill in Section 4. The first step, in Section 9, tests whether a skill helps at all.
1. Goal
Write a few skill files that an AI assistant loads before building or evaluating a pharmacometrics model, so that its work follows the practice already written down in published guidance and in internal guidance documents.
A skill here means a Markdown file of instructions, with a short description that tells the assistant when to load it, in the format Claude Code and similar tools read. One skill per task is the starting assumption: for example one for building a population model, one for evaluating it, one for reporting it.
2. Out of Scope
- Not a new guidance document. The skills restate and cite guidance that exists. Where the sources disagree, the skill links to a position page, as described in the proposal.
- Not a replacement for the modeler’s judgment. A skill sets the checks an assistant runs and reports; the modeler decides what a failed check means.
- No unpublished internal material. Internal guidance enters only as a checklist that has been cleared for publication.
3. Inputs
| # | Input | What it holds | Status |
|---|---|---|---|
| 1 | Reference list | A couple of published sources per topic, the shortest set that covers it | First queue in References |
| 2 | Checklists | Internal guidance documents reduced to checklists that can be published | Not started; needs a clearance route |
| 3 | The skill files | The current version of each skill | First skill in Section 4 |
A first guess at the topics: data assembly, structural model building, covariate modeling, parameter estimation and its diagnostics, model evaluation including visual predictive checks, simulation, and reporting.
4. First Skill: Model Evaluation in nlmixr2
The first skill evaluates a fitted population model in nlmixr2, the open-source R package for nonlinear mixed-effects models. Andy is learning nlmixr2 and wants this skill for his own agentic work, so it has a user from the first day. Model evaluation is also the narrowest topic, and the easiest to test.
Two Modes
| Mode | The assistant sees | The assistant can |
|---|---|---|
| 1. With data | The dataset, the model code and the fit object | Run the fit, make any diagnostic, try alternative models, and report |
| 2. Without data | Exported population-level and run-level outputs, and images of diagnostic plots | Read them, find problems, and suggest next steps for the modeler to run |
Mode 2 is for data that cannot be shown to an AI assistant, which is the usual case for clinical data at a company. It is the harder mode to design, because the skill has to say exactly what to export.
What the Skill Checks
A first list, to be checked against Nguyen et al. 2017 and the nlmixr2 documentation.
- Estimation: did the fit converge, did the covariance step succeed, and are the relative standard errors acceptable.
- Parameters: are the estimates physiologically plausible, and how large are the between-subject variability and the residual error.
- Shrinkage of the random effects, which decides whether the plots of individual estimates in check 6 can be trusted.
- Goodness-of-fit plots: observations against population and individual predictions, and conditional weighted residuals against time and against prediction.
- Visual predictive check, prediction-corrected where dose or covariates differ across subjects.
- Random effects: their distributions, their correlations, and their relationships with covariates.
- Individual fits for a sample of subjects.
What Mode 2 Exports
The run-level outputs (objective function, parameter table with standard errors, shrinkage, run settings) carry no individual data. Plots do: an individual fit shows one patient’s observations, and a goodness-of-fit plot shows every observation as a point. Whether an image of a plot counts as showing the data is a decision for whoever owns the data, and the skill should list the exports in two tiers, summaries only and summaries with plots.
Test Data
Mode 1 can be developed on public data: the example datasets that come with nlmixr2, and synthetic datasets from synpmx. Mode 2 can be tested on the same fits, by exporting their outputs and handing the assistant only those. A test case is a fit with a known problem (a missing compartment, a wrong residual error model) and the finding a good run should report.
5. The Yearly Rebuild
Most of the upkeep could be done by an AI assistant rebuilding each skill once a year from four inputs.
- The existing skill.
- The curated references and checklists in Section 3.
- A new literature review, asking for anything published since the last rebuild that should be included.
- A separate review pass, by a second model or a set of agents, asking what is missing, wrong or out of date in the rebuilt skill.
The rebuild arrives as a pull request: a proposed change on GitHub that shows every line added or removed, reviewed under the rules in the proposal.
6. Risk: Models Outgrowing Skills
Models may soon write their own skills and best practices, well enough that a human-curated skill adds little. Current models already explain a VPC and why shrinkage matters, and the 2026-10-02 search found eleven public libraries of pharmacometrics skills. The parts of this project differ in how exposed they are.
| Part | Exposure | Why |
|---|---|---|
| Skill text on which diagnostics to run and how to read them | High | What a model is most likely to write well unaided |
| Test cases with known problems | Low | They measure any assistant’s model evaluation, whoever wrote the skill |
| Position pages and published checklists from internal guidance | Low | Agreement in the field and unpublished experience are not in the training data |
The test for this risk is a comparison with and without the skill. Run an assistant on the test cases in Section 4 twice, once with the draft skill loaded and once without. If the assistant without the skill already reports the expected findings, the skill text is not needed and the project shrinks to the test cases. If the skill helps, the cases it changes show where.
The design shrinks with the risk. Under the length limit in the proposal, a skill drops whatever models already do well and keeps the contested and the unpublished parts.
7. Who Does the Work
Andy, alone, to start. He is learning nlmixr2 and wants the evaluation skill for his own work, so it is useful to one person from the first day and needs no group to agree on anything. If others find it useful, a person or two may join; a wider group of editors waits until then. A possible outcome is an ACoP poster in 2027 describing the skill and the comparison in Section 6.
8. Open Questions
- Which internal guidance documents exist, and who can clear a checklist drawn from one for publication? This decides whether input 2 in Section 3 is possible at all.
- What does the comparison in Section 6 measure? A count of expected findings reported per test case is the simplest; whether a suggested next step was sensible needs a person to judge.
- Where does the work start? pmxskills.com is registered. ISoP is the right long-term home; the nlmixr2 community could host the first skill, since it already works through GitHub, at the cost of tying that skill to one tool.
- How many skills, split how? By task, as in Section 1, or by model type.
- Where does Mode 2 draw the line on plots? Section 4 proposes two export tiers; someone who owns clinical data has to say which is acceptable.
9. First Steps
- Write three or four test cases with known problems, as in Section 4, on nlmixr2 example datasets or synpmx data.
- Run an assistant on them without any skill, and record what it finds.
- Draft the Mode 1 skill for nlmixr2 model evaluation and run the same cases with it loaded, as in Section 6.
- Define the Mode 2 export and run the same test cases through it.
- Check entries 1, 6 and 9 in References against their sources, since the skill’s checks come from them.