Our continuing struggles with the Monster model

tl;dr. We’re still stuck.

It was 30 years ago that Frédéric Bois, Jiming Jiang, and I published our paper on hierarchical modeling for toxicology. It was a successful paper–it’s been cited over 400 times! It won a major award! More important for me was that working on this project gave me a much better understanding of several different areas of Bayesian inference and computation. I think the greatest impact of this paper has been on informing my thinking, giving me a perspective which has then spread through my research articles, textbooks, and software. I don’t know that it our approach is used much in toxicology, as it took a lot of work just for that one example.

The example is a study of the population toxicokinetics of tetrachloroethylene. Here’s the applied paper with details of the model, and here’s the statistics paper with an explanation of how we did the Bayesian modeling and inference.

The data come from an experiment with 6 people, and for each person there is a multi-compartment model set up as a differential equation with about 15 parameters, and then we had a hierarchical model to allow these parameters to vary by person. We did the computation using Gibbs and Metropolis with a differential equation solver in there too. You can check out the above-linked paper to see how we set up priors and checked the fit of the model to data.

Back in 1993 or 1994 or whenever, it took a lot of work to fit the model. Frédéric programmed it in an C program he wrote called mcsim. But finally it did work, the model fit the data and made reasonable predictions, and we published the papers. As I said, it was a lot of effort for just one toxin, but the work served as proof of concept for the use of Bayesian methods for a complicated model that could not be fit using standard approaches given the sparse data that were available.

I put the example in Bayesian Data Analysis, and then eventually I wanted to fit the model in Stan. When using an example for teaching, it’s good to be able to run the code directly, rather than just pointing to graphs in a published paper.

But it wasn’t so easy to fit the model in Stan! It seems that there are some difficulties with identification of the parameters; perhaps this is one of those cases where computation works ok with a crude algorithm (random-walk Metropolis), but then when you switch to a better algorithm (Hamiltonian Monte Carlo), complexities arise. It’s like picking up a rock and seeing all the worms underneath.

We had some discussion of the problem on the Stan Forums a few years ago. As part of that conversation, Charles Margossian wrote:

It has a lot of moving parts: a stiff ODE that is highly sensitive to parameter values, sparse data, a low number of patients (which makes it it difficult to estimate inter-individual variability), hierarchies across multiple parameters, etc. The workflow is complicated by the fact each model fit takes several hours, even after firing up the cluster and parallelizing within chains. I [Charles] fitted simplifications of the model and got reasonable results. Fixing one or two parameters, or the population variance greatly simplifies the problem. The full model returns divergent transitions and I haven’t found the right parameterization. I tried monolithic centering and non-centering, and things in between. I now suspect the problematic geometry may not be only due to the interaction between population and patient parameters. It’s indeed possible to have a funnel between population parameters, in which case no parameterization can help you. Divergences aside, I don’t trust the estimated parameters, because there are somewhat inconsistent with the priors which were very carefully constructed. Truncating priors, as was done in the original paper, limits our ability to reparameterize the model. When using truncation, I find a lot of the probability mass concentrates at the boundary (note that on the unconstrained scale, you actually have infinite volume near the boundary). All that said, you shouldn’t need to truncate the priors, because the truncation occurs at the tail of the prior distributions. The original paper didn’t use HMC, so they didn’t have access to diagnostics such as divergent transitions. You can also imagine how an accept-reject sampler will interact differently with a truncated distribution. To sum up, I agree this would make a very good case study and writing it has been my ambition for some time. But it’s a challenging model to fit, with behaviors I don’t fully understand (yet). There exist other examples of hierarchical ODE models, built and used with Stan, which can be considered, if the goal is to write a case study. But yes, I think we have a lot to learn from the Monster model.

Lots more in that thread, including this plea from me:

I’d be soooo happy if someone could get the Monster model running in Stan. It’s a great example but it would be so much better if it were runnable. Also it’s buggin the hell out of me that we could it just fine using Metropolis 25 years ago but now we’re getting stuck when using a much better algorithm on computers that are a zillion times faster. So I’d really really really like to get this one straightened out. Anyone who can do this, I promise you a free homemade loaf of bread, an advance look at some blog posts 6 months in the future, and pretty much whatever else you want from me.

We’re still not there yet, and I don’t know the full story. The above offer still stands. Here are the data and some code from Charles.

I’m sharing this example for two reasons:

1. I’d still like to fit the model in Stan. Not just to show off Stan, but because if the model is in Stan, it’s portable. Other people can use it for other, similar, problems. It would be a good teaching example (even though now it’s too late for it to be included in the Bayesian Workflow book). And, given that we did have problems with the computing, it would be good to know what it takes to get the model to work.

2. It’s good to have examples of our fallibility. Of course I’ll share the successes (see here) but the struggles are important too.

P.S. to Bob: I hope you appreciated the careful use of capitalization to convey meaning in the title of this post.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论