Friday, March 6, 2015

Toxicology meets Epidemiology

The Log(dose) Transform

When there are multiple variables that influence, or may influence, a particular health condition, epidemiologists often employ a regression analysis that seeks to tease apart the contribution of each (independent) variable to the outcome (the dependent variable).  When the independent variable is scaled quantitatively (e.g. a dose), the relationship between each variable and the outcome is typically assumed to be linear.  This is partly a legacy from a time when the assumption of linearity was necessary to conduct the analysis at all.  With modern computing power, this is no longer true.  Nonetheless, linearity is usually at least plausible, and therefore a multivariate linear regression is still a useful way to explore potential causal relationships.  However, there are times when the assumption of linearity is not even remotely plausible.  For example, consider the use of log-transformed measurements of a toxicological dose.

After reviewing the available data, the National Academy of Sciences committee charged with evaluating the EPA Reference Dose for methylmercury (NRC, 2000), it was decided that a recently published study conducted in the Faroes Islands (Grandjean et al, 1997) was the most suitable for the development of a Benchmark Dose.  But, when comparing the benchmark dose estimates from different studies, a problem involving model choice was encountered; the three different models considered (K-power, loglinear and square-root) yielded different benchmark dose estimates:

After extensive discussion, the committee concluded that the most reliable and defensible results for the purpose of risk assessment are those based on the K-power model. The argument for this conclusion is as follows. In dose response settings like those with MeHg, when there are no internal controls (i.e., no unexposed individuals) and where the dose response is relatively flat, the data will often be fit equally well by linear, square-root and log models. The models can yield very different results for BMD calculations, however, because these calculations necessitate extrapolating to estimate the mean response at zero exposure level. Both the square-root and the log models take on a supralinear shape at low doses, that is, they postulate a steeper slope at low doses. Thus, they tend to lead to lower estimates of the BMD than linear or K-power models. From a toxicological perspective, the K-power model has greater biological plausibility, because it allows for the dose response to take on a sublinear form, if appropriate. Sublinear models would be appropriate, for instance, in the presence of a threshold. The K-power model is typically fit under the constraint that K ≥ 1, so that supralinear models are ruled out.
Although the committee certainly came to the right conclusion, they should not have needed to think so hard.  Even though they had initially published study results primarily using a loglinear model (i.e. linear regression with a log transformed dose estimate), because of reservations concerning the log transformation, the Faroe Islands study group had already been asked to reanalyze the data with a linear model (i.e. with no log transformation; Budtz-Jørgensen et al, 1999).  But, sometimes a picture really is worth a thousand words.  Consider this modified graph (text and arrows are added)from another neurobehavioral study concerned with arsenic (Wasserman et al, 2004):


Wasserman, et al (2004) is not an unusually bad study.  But, they did have the temerity to show a graph that portrays an absurd underlying mathematical theory – with no data at all. So, here’s the ridicule for them and the many other employers of the log(dose) transform: Toxicological effects get bigger as the dose gets bigger, not as the dose gets smaller.  The log(dose) transform not only flies in the face of toxicology, pharmacology, and biochemistry, it defies common sense.

How Can They Be So Dim?

Good explanations for the continued use of the log(dose) transform in environmental epidemiology may be hard to find, but here is one:
  1. A log(dose) transform may provide a better approximation than a linear one.  Indeed, there are theoretical justifications for this:  Pharmacologicial effects can be supralinear over a narrow dose range.  Nutritional effects can be expected to plateau at high doses.  That said, there is simply no excuse for infinite effects at no dose.


Bad non-mutually exclusive explanations are far more plentiful:
  1. Everybody does it.  The log dose transform has been around for about as long as multivariate regression, and in some fields like neurobehavioral epidemiology, its use is very common.  If everyone in a circle of mutual admirers thinks it’s OK, well then it is.  Except for the other folks.
  2. Statistical Significance.  It may not be plausible, but if it generates an association, that’s what counts.  For publication, anyway. 
  3. The Independent Variables Follow a Lognormal Distribution.  This is a reason that is commonly given, but it makes no sense whatsoever.  The distribution of a set of dose measurements in a given cohort has nothing to do with the shape of the dose-response relationship.  This rationale does blend in with the landscape, unless you look closely, or at all.
  4. Siths Come in Pairs.  Epidemiology is commonly practiced by a Primary Investigator who gathers the data, along with a Statistician who does the analysis.  Whether or not the models used to conduct the analysis are plausible or not may be something neither of them considers.  
  5. They Just Don’t Know Any Better.  Maybe they had a pharmacology or toxicology course once, and maybe they glanced through the chapter on dose-response relationships.  But that was a long time ago and they forgot, apparently.
  6. It’s a Living.  If the papers keep getting published and the grants for low-dose studies keep getting funded, why worry?
Anyway, if you hear a report that an effect increases by some multiple of the dose, regardless of the starting point; try to get the data.

References

Budtz-Jørgensen, E., N. Keiding, and P. Grandjean. 1999. Benchmark Modeling of the Faroese Methylmercury Data. Research Report 99/5. Prepared at the University of Copenhagen, Denmark for the U.S. Environmental Protection Agency.

Grandjean, P., Weihe, P., White, R., Debes, F., Araki, S., Yokoyama, K., Murata, K., Sorensen, N., Dahl, R., Jorgensen, P. (1997).  Cognitive deficit in 7-year-old children with prenatal exposure to methylmercury.  Neurotoxicology and Teratology 19:417-428. 

National Research Council, Committee on the Toxicological Effects of Methylmercury. (2000). Toxicological Effects of Methylmercury. The National Academies Press, Washington, D.C.

Wasserman GA, Liu X, Parvez F, Ahsan H, Factor-Litvak P, van Geen A, Slavkovich V, LoIacono NJ, Cheng Z, Hussain I, Momotaj H, Graziano JH (2004).  Water arsenic exposure and children's intellectual function in Araihazar, Bangladesh.  Environ Health Perspect. 112:1329-33.

Official Post Soundtrack

Man Or Astro-Man? (1999).  Fractionalized Reception of a Scrambled Transmission.  In: EEVIAC: Operational Index and Reference Guide, Including Other Modern Computational Devices, Track 7.

Post Note

Thesis #6.  My wife is a genetic epidemiologist.  I asked her if she would trust the statistician that she works with to pick the model used to analyze the data.  Her immediate response: "No, she doesn't know the underlying biology".  I have a smart wife.

Yes, this post is a bit of a rant.  But, they deserve it.



No comments:

Post a Comment