2  From a question to a type of model

Now you’re thinking: I want to build a model. Where do I start? You go back to reading, but you read differently this time.

Reread for the methods. There’s nothing wrong with rereading a paper you’ve already read; this time you’re paying close attention to how things were done rather than what was concluded. It’s also worth drawing on sources beyond peer-reviewed papers (blogs, books, workshops), because papers tend to explain what was done but not why. The “why” of model development often reveals the tacit knowledge that is needed to make a model work.

Examples include why a researcher chose a particular software or why they used a mathematical shortcut instead of simulation.

These papers you re-read are probably important, so keep them somewhere safe for later reference. We will want to come back to these later to check for disciplinary conventions in writing equations or drawing diagrams. Just pay attention to these conventions now.

Build on a solid foundation, but be novel in just a few ways. Science, like every human activity, is culture-driven. Things get done in certain ways, sometimes for good scientific reasons and sometimes purely by convention, and it’s important to understand what the conventions of your discipline actually are. To move the field forward you want to be novel, but only in a few ways at once. Work that is novel in too many directions, or too far ahead of its time, is hard for others to understand and hard to get through peer review. The sweet spot is the cutting edge, one step ahead of the field. So start with a foundation of what’s already been done, and add to it.

Or notice that the foundation is shaky. Because science is cultural, plenty of things are done simply because that’s how they’ve always been done, and nobody has looked hard at the base. Identifying where a body of work rests on tradition rather than on something robust, and then pulling that part down and rebuilding it on firmer ground, is often a valuable contribution.

Read across disciplines. Some of the best ideas come from cross-pollination. What you’re hunting for are patterns: unexpected similarities in the kinds of problems people face or the designs they use, because a method built for one field can often be transferred to another. A well-known example: ecologists realised their problems looked a lot like the ones social scientists had been solving with Bayesian hierarchical models. Social scientists deal with nested sampling (students within schools within districts), and ecologists have exactly the same structure (quadrats within sites within regions). Because the problems are analogous, the methods transfer, which is why Andrew Gelman’s work has been so influential in ecology even though he’s a political scientist. The borrowing doesn’t have to be between fields as distant as ecology and social science; it can be from a neighbouring corner of your own discipline. Decision-focused ecologists, for instance, have drawn on finance and economics (optimality theory and the like) for environmental decision problems.

Analogues are one of our best tools for generating ideas: find one, and see how far you can push it. Sometimes you get a real insight that way.

Here’s a case study that spans most of the above points. Ecological modellers started borrowing from machine learning to use methods with great predictive credibility for ecological problems. This turned out to be a fruitful area of research, particularly for the topic of predicting where species live. However, this new culture of applying machine learning to ecological data was built on some shaky causal foundations. The issue is that what works well for prediction doesn’t always work well to infer causes. For good prediction, we just need to find correlates of the outcome we’re interested in. But as everyone knows, correlation does not always imply causation (more on when it does and doesn’t in Shipley 2016).

The predictive modellers started over-interpreting their machine learning models, taking good predictors to be causes of why species live where they do. One student addressed multiple of these problems in her PhD thesis. Firstly, she showed that interpreting good prediction as also meaning cause is bad logic (Arif and MacNeil 2022). Then she rebuilt the shaky foundations by cross-pollinating methods from the socio-economic sciences, specifically structural causal modelling (Arif and MacNeil 2023). So she brought new methods to address the gap that she had created by tearing down part of the foundations around using prediction in ecological causal inference.

Arif, Suchinta, and M. Aaron MacNeil. 2022. “Predictive Models Aren’t for Causal Inference.” Ecology Letters 25: 1741–45. https://doi.org/10.1111/ele.14033.
———. 2023. “Applying the Structural Causal Model Framework for Observational Causal Inference in Ecology.” Ecological Monographs 93 (1): e1554. https://doi.org/10.1002/ecm.1554.
Shipley, Bill. 2016. Cause and Correlation in Biology: A User’s Guide to Path Analysis, Structural Equations and Causal Inference with r. 2nd ed. Cambridge University Press. https://doi.org/10.1017/CBO9781139979573.