Wednesday, May 11, 2016

Biological Problems

The Perils of Multivariate Linear Regression

As is often the case in the epidemiological literature on environmental influences on neurobehavioral development, Bowers and Beck (2006) noted that a paper by Lanphear et al (2005) “has suggested the existence of a supra-linear dose–response relationship between environmental measures such as blood lead concentrations and IQ”.  They then produced an analysis that indicates that the apparent supralinearity is an artifact resulting from the way the data were analyzed.  They stated their conclusion as follows:
Results of the analyses show that a supra-linear slope is a required outcome of correlations between data distributions where one is lognormally distributed and the other is normally distributed. 
While their mathematical analysis was indubitably correct, the way Bowers and Beck reported the results left something to be desired.  How the data are distributed is not really the issue at all.  Instead, the mathematical artifact they found results from conducting linear regression analyses with log transformed data.  If data from a normal distribution, or any other distribution, were log transformed prior to the regression analysis, then the same result would be obtained.  Furthermore, as demonstrated by Jusko et al (2006), a linear regression without log transformation with data drawn from a lognormal distribution does not result in a supralinear dose-response relationship.

To their credit, Hortung et al (2006) also understood that the real issue is the shape of the dose response relationship rather than the distributions that either the dependent or independent variables follow.  They therefore protested that Lanphear et al (2005) had considered the likely shape of the curve before conducting the regression analysis:
The shape of the exposure–response relationship was determined to be nonlinear insofar as the quadratic and cubic terms for concurrent blood lead were statistically significant (p < 0.001 and p = and 0.003, respectively).  Because the restrictive cubic spline indicated that a log-linear model provided a good fit to the data, we used the log of concurrent blood lead in all subsequent analyses of the pooled data.
But, there are many problems with this justification.  First, it is not at all clear how a spline analysis specifically supports a log-linear model, as opposed to other potential nonlinear models (e.g. a Hill function).  Second, there was no consideration of biological plausibility.  Like Bowers and Beck (2006), Lanphear et al (2005) seem to think establishing a causal relationship is a mathematical problem rather than a biological one.  Third, they did not consider the possibility that other covariates might explain the apparent nonlinearity.   Yet, off they went, and a dose-response model that predicts infinite large effects as the dose approaches zero was the inevitable result.  For all practical purposes, Bowers and Beck (2006) were entirely correct.  

Besides the fact that a loglinear dose-response model is a very poor theory, there is a more general lesson to be learned:  A multivariate regression with assumed quantitative relationships between the variables being modeled is highly prone to error.  While a loglinear relationship is obviously wrong, a linear relationship isn’t necessarily right either.  Correlations between variables may result in attribution of mismodeled causal effects to a variable that has no causal effect at all.  For example, if the relationship between socioeconomic status (i.e. the HOME score) and IQ is nonlinear with bigger impacts with low scores and negatively correlated with exposure to an environmental chemical, the some of the socioeconomic effect will erroneously appear to be a low dose effect attributable to the environmental chemical.  There are many other possible explanations as well, all of whicih are more probable than a dose response model than predicts incremental effects to get bigger as the dose gets smaller.

Process vs. Theory

Biological complexity often makes the pronunciation of definitive truths doe Medicine and Public Health practically impossible.  While relying on expert opinion is a common solution to that problem, that solution does not work well when opinion is divided.  As a means of coping with that problem, institutional decision making processes often employ structured evaluation systems to sort through what can often be a voluminous set of scientific literature.  The Safety Assessment methodology that is typically used for premarket approval evaluations is an example.   This description of Evidence-Based Medicine conveys the general ethos of such efforts:
Whether applied to medical education, decisions about individuals, guidelines and policies applied to populations, or administration of health services in general, evidence-based medicine advocates that to the greatest extent possible, decisions and policies should be based on evidence, not just the beliefs of practitioners, experts, or administrators. It thus tries to assure that a clinician's opinion, which may be limited by knowledge gaps or biases, is supplemented with all available knowledge from the scientific literature so that best practice can be determined and applied. It promotes the use of formal, explicit methods to analyze evidence and makes it available to decision makers.
There are two key concepts at work here.  First, the “beliefs of practitioners, experts, or administrators” are getting kicked to the curb in favor of “evidence”.  If you thought the beliefs of experts were based on scientific evidence, then you were misinformed, apparently.  Secondly, there is an emphasis on the “use of formal, explicit methods”, which also serve to limit subjective influences on the evaluation process. 

Experts are not always trustworthy, so the desire for a transparent process is entirely understandable.  But, getting a trustworthy process to replace the experts is easier said than done.  The process has to be designed by somebody, and that usually means experts.  There is also apt to be a negotiation process involved in getting the process to be accepted, so subjectivity isn’t really completely avoided.   But perhaps the bigger problem is that trying to deal with complex biological issues with a formula may often be rather stupid.  If all the studies show the same result, then it really isn’t going to matter whether the decision making process is expert-based or evidence-based.  If the results are different, then the systematic review may succeed at identifying the higher quality studies and grading the general result.  But, it won’t explain why the results are different.  It won’t figure out why a treatment may works sometimes, but not others.  That will take biological theories, and like the biases the evidence-based systems strive to avoid, those are subjective.   There are likely to be different theories, of course, and then the experts will inevitably get into a debate over which are more likely.  But, guess what, that’s the way science works: Trying to eliminate all potential bias with formulae will also eliminate scientific progress.

By all means, more transparency is needed.  In particular, let’s not trust authors to have the last word on how the data they have collected are analyzed and published.  Medical researchers and epidemiologists are notorious for not sharing original data involving human subjects, even when they are legally required to do so (Panhuis etal, 2014; Longo and Drazen, 2016).  That will allow better theories to flourish, and poor theories to flounder.

References

Bowers TS and Beck BD (2006).  What is the meaning of non-linear dose-response relationships between blood lead concentrations and IQ?  Neurotoxicology 27:520-4.

Hornung R, Lanphear B, Dietrich K. (2006).  Response to: “What is the meaning of non-linear dose–response relationships between blood lead concentrations and IQ?”.  Neurotoxicology 27:635

Jusko TA, Lockhart DW, Sampson PD, Henderson CR Jr., and Canfield RL (2006).  Response to: “What is the meaning of non-linear dose–response relationships between blood lead concentrations and IQ?”.  Neurotoxicology 27:1123–1125.

Lanphear BP, Hornung R, Khoury J, Yolton K, Baghurst P, Bellinger DC, Canfield RL , Dietrich KN, Bornschein R, Greene T, Rothenberg SJ,8, Needleman HL, Schnaas L, Wasserman G, Graziano J,13 and  Roberts R. (2005).   Low-Level Environmental Lead Exposure and Children’s Intellectual Function: An International Pooled Analysis.  Environ Health Perspect. 113: 894–899.

Longo DL and Drazen JM (2016).  Data Sharing.  N Engl J Med 374:276-277.

Panhuis WG van, Paul P, Emerson C, Grefenstette J, Wilder R, Herbst AJ, Heymann D, and Burke DS (2014).  A systematic review of barriers to data sharing in public health.  BMC Public Health 14:1144.

Official Post Soundtrack


Jackson, J (1980).  Biology.  In: Beat Crazy, Track 9.

Post Notes

Thesis Post #65.  This covers some of the same ground as Toxicology Meets Epidemiology, but with a more philosophical overview. 

Thursday, May 5, 2016

Mixed Probability Calculations

Probability is the Guide of Life

For personal decisions, theoretical uncertainty is the far more familiar form of probability.  If two different sources of information lead to different courses of action, then you have to either decide who and what to trust or hedge your bets.  However, the probability of chance that is amenable to a mathematical treatment and is the main form found in academic discourse can be important too.   The relative importance of the two probabilities can vary with the problem.  Sometimes one or the other will dominate, while in other instances both are important.  Recreational betting games serve as an example:
  • Roulette.  Betting on a roulette wheel is purely a game of chance.  The odds and a long term expected return can be calculated very accurately.  Well, unless the game is fixed.
  • Horse Racing.  In theory, some horses are faster than others – chance has very little to do with who wins.  Sure, historical records are important, but that’s mainly because they indicate which horses are fast and which ones are not.
  • Poker.  The odds that a certain card or cards will turn up can be calculated, and the game of poker can be simply played as a game of chance.  But good poker players also take the mannerisms of their opponents into account when they bet, which turns poker into a mixed probability game.

It’s the last category of problems that make risk analysis interesting.
 

Betting on the Single Instance

If you are betting on a single instance (i.e. what to do now), then boiling down theoretical probability and statistical probability into a single judgment or number is essential.  A simple equation will suffice to represent this notion:

pTotal = pTheory * pChance

If the roulette wheel is fair, then pTheory =1, and therefore pTotal is dominated by calculating the odds.  If the fastest horse wins, then pChance is 1, and pTotal is dominated by pTheory.  When gamblers bet on a horse, converting horse theory to a numerical value is exactly what they do.  Poker players have a tougher calculation – not only do they have to know the odds of a card turning up, they also have to assign a probability to the notion that bluffing will work, or that their opponent is bluffing.

Betting on the Series

But once the bet becomes about the long run, or about public health instead of an individual, then the calculation is quite different.  It’s a two dimensional problem where the primary goal is to predict the frequency of a result or different results, and there will also be uncertainty about estimated frequencies.  The probability calculation isn’t the same any more.  The probability of chance is often a statistical frequency instead.  In fact, it may or may not be a theoretical frequency.  For example, there can be a range of statistical estimates that range from purely empirical to purely theoretical.  An historical record with a large number of observations may justify a frequency estimate with no theoretical uncertainty.  On the other hand, a fewer number of observations may serve to support a statistical theory instead, which begets theoretical uncertainty.  The frequency calculation is now a function instead of a single number, so the relationship between theoretical probability and the frequency of occurrence is now something far more complicated:

p(Frequency) = pTheory(pChance)

Empirical observations may also be used to disprove a theory too.  For example, a large number of observations may show a particular die to be unfair.  The again, there may only be enough data the favor one theory over another without being able to conclusively decide that one is indubitably correct.  That means you are going to need a probability tree

Quantifying Theoretical Probabilities

Frequentist probability schemes tend to acknowledge theoretical uncertainty (e.g. as “systematic error”), but then go on to ignore it.  On the other hand, Bayesian probability schemes typically treat theoretical and statistical probabilities interchangeably.  If you are betting on a single instance, that works reasonably well.  Updating a theoretical prior with data can gradually transform the probability into one of chance – the more data there are, the less the theory matters.  But it isn’t really very scientific.  If they were used to discriminate among alternative theories, the data might be put to better use.  That problem is even more critical for the estimation of long run frequencies.  Updating the parameter estimates for a model that has been proven to be wrong doesn’t make much sense.

Since it really is more consistent with how scientific knowledge is developed, explicitly assigning probabilities to theories is a better strategy for long-term issues where knowledge may be expected to progress.  Since theoretical probabilities are inherently subjective, it is hard to improve upon convening a panel of experts to weigh the scientific evidence.  Even if the experts don’t get it quite right, or they aren’t the right sort of experts, the process of assigning probabilities to competing theories creates an occasion for scientific discussion.   As long as no one thinks that probabilities assigned to theories are the gospel truth, it’s all good in my book.

As a recent example, Trasande et al (2015) provided an overview of the efforts to characterize the theoretical probabilities for causal theories involving potential health effects of Endocrine Disrupting Chemicals (EDCs):
We now describe the general methods used to attribute disease and disability to EDCs, to weigh the probability of causation based upon the available evidence, and to translate attributable disease burden into costs. During a 2-day workshop in April 2014, five expert panels identified conditions where the evidence is strongest for causation and developed ranges for fractions of disease burden that can be attributed to EDCs.
I have more than a few quibbles with exactly what they did, ranging from how the problems were characterized in the first place (i.e. by presuming independent attributable risks), the use of implausible dose-response models, the lack of serious consideration of other (i.e. non-EDC) causal factors, and the relationship between association and causation is all-or-none.   Also, because the probability assignments are subjective, a two-day workshop of experts with similar interests is not really sufficient for a decision involving the economic impacts that are alleged, so I don’t recommend taking these estimate as the last word. However, praise for the process is well deserved.  Nonetheless, as it pertains to the present topic of discussion, there is one error in how the theoretical probability was employed after it was arrived at that must not go unnoticed:
Finally, recognizing that attributable cost estimates were accompanied by a probability, we performed a series of Monte Carlo simulations to produce ranges of probable costs across all the exposure-outcome relationships, assuming independence of each probabilistic event. Separate random number generation events were used to assign 1) causation or not causation, and 2) cost given causation, using the base case estimate as well as the range of sensitivity analytic inputs produced by the expert panel. To illustrate with an example, for an exposure-outcome relationship with an 80% probability of causation, random values between 0 and 1 in each simulation led to the first step, which either assigned no costs (random value ≤ 0.2) and costs (random value > 0.2).
If the problem required the combination of both theoretical and statistical probabilities, the use of the probability tree in a Monte-Carlo simulation would be appropriate.  However, there is a problem in implementation that arises from the fact that a causal probability is NOT a probability of chance: A theory is either true all the time or false all the time, and the entire cost estimate is dependent (so, no you can’t assume independence) on the truth of the theory.  So, using a causal probability to calculate the probability of an event is inappropriate.  Instead, the logic should go like this: Since all of the causal probabilities have a probability of less than 95%, the lower bound cost estimate of all of the end points should be zero (see table four).  For those endpoints with a causal probability of less than 50%, the central estimate should be zero as well.   

Reference

Trasande L, Zoeller RT, Hass U, Kortenkamp A, Grandjean P, Myers JP, DiGangi J, Bellanger M, Hauser R, Legler J, Skakkebaek NE, and Heindel JJ (2015).   Estimating Burden and Disease Costs of Exposure to Endocrine-Disrupting Chemicals in the European Union.  J Clin Endocrinol Metab 100: 1245–1255.

Official Post Soundtrack


Cars, The (1978).  All Mixed Up.  In: The Cars, Track 9.

Post Notes

Thesis Post #64.  If someone can figure out a way to short their bet on all those IQ points, I'm all in.

Wednesday, April 27, 2016

Individual Choice

Public Health Value Judgments

When you can’t have it all, which is pretty much all the time, it is necessary to set priorities.  Other people (e.g. friends, relations, employers, and the government) are often there to help you set your priorities – whether you want them to or not.  But still, everybody does have to make their own choices on occasion.   For example, unless your mom, spouse, or religion chime in, you can choose what fish you eat and how much all by yourself.  You probably already know whether you like fish or not, and you also probably know how much it costs.  If there are other factors that go into the decision, then you will need to know what they are.

Unfortunately, the trend in public health these days is give consumers food consumption advice without exactly saying why.  There is no good reason for this that I know of, but I am aware of two of the bad ones.  First, not getting in to the gritty details avoids political controversy stemming from scientific uncertainties.  That doesn’t mean the advice in necessarily bad, but then again maybe it is.  Second, doling out public advice can be a career all by itself, and career advisers often care more about protecting their jobs than whether or not the advice is sensible.  So, at best, food consumption advice is an expression of the social values of the people giving advice - which may or may not correspond to your values.  At worst, the advice doesn’t reflect anyone’s valuation at all.  As a result, distrusting public health advice is generally a pretty good idea, especially when the advice isn’t accompanied by some intelligible reasons for it, which will also permit you decide for yourself if those reasons are good enough for you.

Grading on the Curve

While psychology studies are sometimes grounded in physiology with physical measurements (e.g. nerve conduction velocity), most epidemiological studies concerned with neurobehavioral development largely employ batteries of tests that reflect the social science interface of psychology.  Since the value of these tests is subjective, they are standardized by determining how subjects “normally” perform.  Since variation in performance on tests is normal, the tests are typically given numerical values that reflect how far above or below average a score is relative to much it normally varies.  There are a two major problems with this.  First, how much a test score varies isn’t necessarily a good indicator of how much performance on the test really matters.  Second, defining what a “normal” population can be rather arbitrary.  For example, the Denver Developmental Screening Test was originally standardized in Denver, while the Boston Naming Test was developed in Boston.  Yet, what is normal in Denver may be somewhat different from what is normal in Boston.   For a book length discussion of the problems with standardized testing, see Gould (1981).  Nonetheless, standardized psychological test batteries are more objective and reproducible than a doctor’s or a teacher’s opinion, and they are widely used for that reason.

The most basic normalized scale used for psychological testing is the Z-score where the difference between the average and test score is divided by the standard deviation.  As a result, a -1 signifies a test result that is one standard deviation below the average, while a value of +1 signifies a test result that is one standard deviation above average.  Other standardized tests are often scaled with modified with modified Z-scores.  In particular, the Intelligence Quotient (IQ) is scaled with a mean is defined to be 100 and the standard deviation is 15, while Scholastic Aptitude Tests (SAT) have a mean of 500 and a standard deviation of 100.  The following table compares how test results are scaled with each method:


2 SD below
1 SD below
Average
1 SD above
2 SD above
Z-Score
-2
-1
0
1
2
IQ
70
85
100
115
130
SAT
300
400
500
600
700

Individual Neurobehavioral Risks and Benefits Arising From Fish Consumption

So, let’s talk about fish.  The point of the preceding discussion is that there is evidence that the consumption of fish during pregnancy may have both bad (from methylmercury) and good (from omega-3 fatty acids or perhaps something else) effects on future neurobehavioral performance of the child – and the effects aren’t exactly the same.  Plus, I’m not going to tell you what you should do.  I am leaving that to be your problem, and since it really isn’t a simple decision I have no idea what you will decide.  However, I will do my best to supply some reliable information. 

My main vehicle for information delivery is an Excel-based program, which is a slightly modified version of a risk-benefit assessment model that I developed while I was at the FDA (2014).  Although there are some other minor modifications as well, the main difference is that this version of the model is intended to estimate risks for a specific individual.  However, if you don’t have Excel, or it is more trouble than it is worth, here are some sample results that give a feel for what the program does:
  • Consuming a high mercury such as swordfish fish twice a week during pregnancy will result in a developmental delay of the age at which a toddler learns to walk of about a week.  The uncertainty associated with this estimate ranges from 0 to about three weeks.
  • Consuming one can of albacore tuna and one can of albacore tuna once a week during pregnancy will result in an increase of about 2.5 IQ points in the child.  However, there may be a decrement in IQ of about 0.3 points, or the increase may be as much as 3.5 points.
  • Consuming salmon once a week will result in a projected increase in performance on the Verbal SAT of about 25 points.  However, given the many uncertainties association wit the estimate, the increase may be as little as 0 or as much as 35 points.

Software

This Excel macro, Personal_Seafood_Net_Effect_Estimator.xlsintegrates four components presented earlier:



References

Gould, SJ (1981).  The Mismeasure of Man.  W. W. Norton & Company.


Official Post Soundtrack

Ponty, J-L (1983). Individual Choice.  In: Individual Choice, Track 5.

Post Notes

Thesis Post #63, and the fifth of a series of five.

Sunday, April 17, 2016

Dear Journal Editors

Rejected

Thank you very much for reviewing my manuscript entitled “Plausible In, Plausible Out: A Bootstrap Methodology to Characterize the Uncertainty Associated With Dose-Response Modeling”.  I am disappointed in the result, of course.  However, as it met the same fate as every other paper on model uncertainty that I have sent to the various and sundry editors that the Journal has had over the last 25 years, I am not terribly surprised.

I think I will not attempt to rewrite the paper to make it more acceptable.  The paper says what I want it to say and I don’t think any of the major suggestions made in the review will make it any better.  A similar example to that discussed in the manuscript is also in the FDA (2016) assessment on arsenic in rice released two weeks ago (see section 9.4), and I suppose my purposes will be better served by working the rest of the text into my other writing projects.  Nonetheless, for the benefit of my colleagues who are interested in model uncertainty in general and this paper in particular, I would like to address some of the comments made in the course of the review.

A Methodology Paper

The paper I submitted is a discussion of two methodological developments used in USFDA assessments for arsenic in apple juice (Carrington et al, 2013) and rice (USFDA, 2016).  The first method involves the use of a parametric bootstrap simulation to propagate the uncertainty associated with dose estimation into the characterization of the dose-response relationship.  Putting error bars on the doses is a novel technique and as just about everyone I know thinks it is a pretty good idea, this was the impetus for writing the paper in the first place.   Although the reviewers didn’t seem to think this technique was remarkable, at least they didn’t object.  I guess it really is a pretty obvious thing to do once you’ve thought of it.  The second reviewer did suggest that additional details be provided about the input distributions for the dose estimates.  I decided not to do that when I wrote the paper because I thought getting into specifics would be a distraction from the more general idea of allowing dosimetric uncertainties to be represented in a dose-response analysis.   If I get around to reworking the analysis for some other purpose, I will heed those suggestions to the extent that I can; many of the issues raised by the second reviewer resulted from the necessity of working with published summary data rather than observations from individual subjects.  However, for a methodological presentation where no importance is attached to the actual results at all, I don't think any of that matters.

Against the wishes of potential FDA coauthors, I also chose to include a discussion of model uncertainty.  Since this was and is the hot button political topic for any risk assessment involving arsenic and any other chemical hazard worthy of attention I thought it would be a serious omission to not include it, especially since there may often be an interaction between dosimetry and empirical weighting of alternative models.  As near as I can tell, the reviewers haven’t raised any serious objections to anything I said about model uncertainty, but it is very clear that they really don’t like the way I said it.  This may be partly due to the fact that the reviewers didn't take the discussion in the methodological context that was intended, but since this seems to be the reaction I always get when I try to talk about model uncertainty, I think there is more to it than that.  I think I understand the nature of this editorial issue far better than I used to, so I will take this opportunity to explain.

Unfinished Science

Model uncertainty is subjective.  Some people have it and others don’t.  I think the two basic causative factors are as follows:
  • The model has to be thought of as a theory.  Even if it is only approximately correct, the mathematical model has to convey some truth that is not evident from isolated observations.
  • There has to be more than one model-theory.  If only one model is under consideration, then there is no uncertainty. 

Taken together, these two criteria basically mean that if you are afflicted with model uncertainty then you are thinking like a scientist.  But here’s the thing; if you have it then you can’t really talk or write about model uncertainty in the objective third person writing style generally preferred by governments and journal editors.  You can describe a model uncertainty as a psychological phenomenon as I just have, but that is pretty much the end of the third person road.  If you think it is just me who can’t write properly, go back and read the most widely cited paper on model uncertainty ever written (Hill, 1966): It is written almost entirely in the first or second person.  For example, the problem the paper sets out to solve is stated as follows:
Our observations reveal an association between two variables, perfectly clear-cut and beyond what we would care to attribute to the play of chance. What aspects of that association should we especially consider before deciding that the most likely interpretation of it is causation?
Part of the problem is that scientists don’t write their papers in the same manner as they converse in private.  When papers are written, it is often because model uncertainties have been resolved and what were once just theories are reported to be objective realities.   But risk analysts can’t do that.  There are many model uncertainties that have not been resolved and they may never be, which leaves us with probability trees and subjective weight-of-the-evidence evaluations.

My Problem

The apple juice and rice risk assessments both used probability trees to depoliticize arguments over which dose-response model “should” be used to characterize the causal relationship between inorganic arsenic and cancer.  I am happy to report that this strategy worked as I hoped that it would.  The two assessments employed somewhat different strategies for assigning probabilities to alternative models.  Because I think it is more consistent with how scientists actually think, I prefer the strategy used in the rice assessment that largely relies on expert opinion. 

My only reservation is that the probabilities used for the rice assessment relied only on my opinion.  As an expert, assigning probabilities to theories was implicitly part my job description at the FDA, and I’m not complaining about having to do what I was paid to do.  I have a PhD in Pharmacology and long experience in modeling dose-response relationships, so I don’t feel unqualified.   However, my primary area of expertise is neurotoxicology rather than cancer biology, and I am far more familiar with the literature on lead and methylmercury than arsenic.  So, it would have been nice to have other experts involved, especially if the stakes are raised from just setting guidance values for apple juice and infant cereal that have relatively little economic impact to suggesting that consumers modify their rice intake.  But in order move from the subjective “I” to the intersubjective “We”, I think that those of us who are afflicted with model uncertainty need to be permitted to write in the same way as we converse among ourselves, and we can’t do that if we are forced to pretend to objectivity that we really don’t have.

References


Carrington CD, Murray C, and Tao, S. (2013). A Quantitative Assessment of Inorganic Arsenic in Apple Juice

Hill, Sir Arthur Bradford (1965).  The Environment and Disease: Association or Causation?  Proc Royal Soc Med 58:295-300.

U.S. Food and Drug Administration (2016).  Arsenic in Rice and Rice Products Risk Assessment.

Official Post Soundtrack


Green Day (1997).  Reject.  In: Nimrod, Track 14.

Post Notes

Thesis Post #64.  Even though I will probably group it in the arsenic series, this post is really about academic politics.

Tuesday, April 5, 2016

Arsenic in Rice: Another Bloody Election

Toxic Endpoint Election

The basic idea of the risk assessment paradigm is that it has two steps (NRC, 1983).  In the first step, the risk assessors produce the best information they can about a particular hazard that they have identified.  In the second step, the risk managers take that information and do their best to manage the risk for the benefit of the public.  But the risk assessment paradigm has an evil twin.  The risk propaganda paradigm begins with identification of an issue that the managers would like to be in control of for the benefit of themselves and their friends.  The initial process is followed by a risk caressment process, where the analysts produce a result that justifies a decision that has already been made.  Perhaps the most well-known example of the risk propaganda paradigm is the GW Bush White House using weapons of mass destruction (WMD) as a pretext for the invasion of Iraq

Unlike the apple juice risk assessment, the FDA (2016a) risk assessment for arsenic in rice contains a section on noncancer endpoints.  The National Institute on Environmental Health Sciences (NIEHS) has been funding many “low-dose” epidemiology studies over the last 15 years, there are many new reports in the literature of associations between arsenic and many different toxic endpoints.  However, since there has been an emphasis on studies on women’s and children’s health, that is what most of the resulting publications have been concerned with.  The problem is that the modus operandi in epidemiology these days is to conduct a post-hoc analysis that yields a statistically significant result, write a paper, and then call up the university press office to report the association to the public.  But it’s not an association or statistical significance that matters, it’s causality.  And most environmental epidemiology studies these days don’t seem to be designed or analyzed with the goal of demonstrating causality.

Fortunately, the National Academy of Sciences (2013) recently reviewed the literature on arsenic and sorted the various potential endpoints into “tiers” that grouped them based on the strength or weight of the evidence for a causal relationship.  The noncancer endpoint that fell into the top tier was cardiovascular disease.  The reason for that is pretty simple; while there are multiple studies showing a dose-response relationship between inorganic arsenic and cardiovascular disease, there is little to be found beyond statistical significance for the other endpoints.  So, if you want a quantitative assessment for a noncancer endpoint, cardiovascular disease is pretty much the only choice.

But, that’s not what the FDA did.  They skipped over cardiovascular disease and went for the Tier 2 and Tier 3 endpoints associated with pregnancy and childhood development.  The fact that there were no data to support a quantitative assessment was found to be lamentable, but it left them undeterred.  That might seem unfathomable, but once you understand that it’s the risk propaganda paradigm at work, some reasons for it aren’t that hard to come up with: a) NIEHS wants to justify funding studies that aren’t really very useful, b) the Society of Toxicology wants to create jobs for women who understand pregnancy and children so much better than men do, c) issues concerned with environmental effects on pregnant women and children are part of the Democratic party platform, and d) all of the above.  I think d is the correct answer, but it's mostly a.

Risk Manglement

In its purest form, the risk propaganda paradigm isn’t used very often.  That is because it often ends up with lies that are obviously not true.  Propaganda works much better when it is at least partially true.  The WMD ruse is a pretty good example.  Iraq was invaded, but as it turned out, the WMDs just weren’t there.  The noncancer risk assessment for arsenic in rice is like that too.  The table of contents suggests that there is a risk assessment, but if you look inside it’s just not there.  No quantitative risk assessment, no safety assessment, just an exposure assessment.

Yet the guidance issued by the FDA (2016b) relies primarily on the noncancer nonassessment anyway.  Perhaps that is because the quantitative assessment indicates that, as with apple juice and as usual, setting levels are not a very effective way to reduce the risk from naturally occurring contaminants.  If your management strategy isn’t going to work, then maybe it is better to use a nonassessment that doesn’t show anything at all.  But, the cancer assessment really is a better justification.  The evidence for an effect is stronger, there is some basis for quantifying how big the effect might be, and there is substantial evidence that exposure earlier in life is more important.  The FDA survey indicated that levels of arsenic in infant cereal are actually higher than rice in general, and there is no good reason for that.  A “37 percent reduction in lifetime cancer risk attributable to brown-rice infant cereal consumption” isn’t much, but it’s something.

In addition to proving a guidance level for inorganic arsenic in infant cereal, the FDA (2016c) also suggests that infants and pregnant women should modify their rice consumption.  Again, the cancer risk assessment is arguably sufficient justification for not using rice cereal as the major staple for infants.  But, the advice to pregnant women rests solely on the dubious and unquantifiable Tier 2 and Tier 3 effects.  The review quoted by the FDA as justification for that focus also lists cadmium, copper, iron, manganese, and zinc as potential concerns in addition to arsenic (Wright and Bocarelli, 2007).  What about them?  If simply declaring a causal relationship without any consideration of the relationship between dose and health outcome is sufficient for consumer advice, then what are pregnant women supposed to eat?  Acrylamide – well there goes the bread.  Water is out too – hyponatremia can cause severe brain damage.  Perhaps they can still eat cake.

But, never mind the infants and pregnant women, what about the rest of us?  The guidance document says that

The FDA did not find a scientific or public health basis to recommend that the general population of consumers change its rice consumption based on the presence of arsenic.

What the hell happened to the Tier 1 effects?  Lifetime exposure to arsenic is still causing lung and bladder cancer, right?  What about those cardiovascular effects that the FDA skipped over?  What about the fact that adult males are exposed to about twice as much arsenic from rice products (it’s beer) as adult females?  How come we don’t get advice too?  I’m going to drink the beer anyway, but still I’d like to feel needed.



I’d just blame the Center for Food Safety and Applied Nutrition for being so stupid, but I also know that they bear only proximal responsibility for this gross insult to the intelligence of every woman, man, and child in the United States.  I figure NIEHS, the rest of HHS, the EPA, the Society of Toxicology, the Democratic party, and the bankers who supplied campaign funds for a voting bloc also helped set the propaganda campaign in motion. 

References

National Research Council (1983).  Risk Assessment in the Federal Government: Managing the Process. National Academy Press, Washington, DC.
U.S. Food and Drug Administration (2016a).  Arsenic in Rice and Rice Products Risk Assessment.
Wright, RO and Baccarelli A. (2007). Metals and neurotoxicology. The Journal of Nutrition 137: 2809–2813.

Official Post Soundtrack

Killing Joke (1996).  Another Bloody Election.  In Democracy, Track 10.

Post Notes

Thesis Post #61 and part two of two on the FDA Arsenic in Rice Risk Assessment.  You may notice that in the previous post I refer to my former agency as "we", but as "them" in this one.  That's because the agency I used to belong to doesn't exist anymore.


Monday, April 4, 2016

Arsenic in Rice: Just One Victory

Breaking the Mold

When arsenic in apple juice became a public issue in 2011, the FDA needed a risk assessment.  It was generally presumed by agency management that the toxicologists who worked for the Center for Food Safety and Applied Nutrition should produce a proper risk assessment; meaning an assessment that conforms to standardized EPA guidelines.  But, there was a problem: The most recent EPA (2005) cancer risk assessment guidelines don’t even begin to work for a naturally occurring toxicant like arsenic.  If they did, the EPA wouldn't still have a cancer slope factor that hasn’t been updated since 1988.  If they did, the EPA wouldn't have needed to contract out the cost-benefit analysis for the 2001 Arsenic Drinking Water Rule.  If they did, the FDA wouldn’t be in the position they were in.

So, for arsenic in apple juice we did something else instead (Carrington et al, 2013).  A similar analysis was released last week on the topic of arsenic in rice (FDA, 2016).  While the dose response analysis in those assessments doesn’t conform to the 2005 EPA guidelines, it is consistent with the general notion of separate risk assessment and risk management processes (i.e. NRC, 1983), and it is also consistent with the less prescriptive 1986 EPA cancer risk assessment guidelines.  The most significant departure of the FDA analyses from those guidelines is that they use a probability tree (another old idea) to characterize the uncertainty with the extrapolation from high to low doses.  The most obvious result of using this technique is that it does a more comprehensive job of characterizing the uncertainty associated with generating low-dose estimates from high-dose observations.  But, the more important advantage is that it changes the choice of what dose-response model “should” be used for making a regulatory decision into a different issue from its historical counterparts:
  • NRC 1983 and 1986 EPA guidelines.  A default model was justified by policy unless a scientific argument could be presented for deviating from it.  The problem with this was that toxicologists could never make a scientific argument that was certain enough for the policy default to be overturned
  • 2005 EPA Guidelines. The policy default became mandatory.  Scientific arguments were prohibited by the Point of Departure.
  • Arsenic in Rice Assessment.   While the use of a probability tree means absolute certainty is not required, scientific arguments are mandatory.  Using a probability tree doesn’t eliminate the possibility of political bias, at least it doesn;t require it.  Furthermore, probability trees at least make it possible to frown upon self-interested biases as they occur.

The subjective weights and probabilities in the rice risk assessment are my own.  While no one frowned upon my potential political biases, no one who reviewed the risk assessment expressed an opinion about how the alternative models were weighted either, nor did they suggest alternative models that might be used instead.  That’s a shame that I attribute to the fact that the 2005 guidelines essentially shut down the market for theoretical reasoning in toxicology.  Instead, the current fashion seems to be that a statistical analysis will resolve all the uncertainties with a purely empirical approach – which means Toxicology has largely been supplanted by Epidemiology.  I beg to differ.

Another, more novel, technique introduced in the apple juice and rice assessments is that both analyses use a parametric bootstrapping technique (akin to a Monte-Carlo simulation) to represent the uncertainties in the dose estimates used to characterize the dose-response relationship.  As with the probability tree, there are two advantages.  First, there is better characterization of the uncertainty associated with the dose-response relationship.  Second, (and again, more importantly) it is no longer necessary to decide when the dose estimates for human epidemiology studies are “good enough” to proceed with a characterization of a dose-response relationship. 

The Devils in the Detail

Using a probability tree to raise the lid on Toxicology reveals a wide array of wriggling quantitative issues.  One of the reasons the EPA guidelines have proscribed relatively simple dose response models is because choosing a model for political reasons (e.g. to be precautionary) is only possible when there is a clear relationship between a scientific assumption made and its regulatory implications.   With a simple model that choice can often be made irrespective of any other scientific issues.  With more complicated models that may not be true.  For example, whether or not a nonlinear dose-response model is “more conservative” (i.e. yields a higher risk estimate) may depend on what the dose is.

The impetus for precautionary assumptions has led to the notion that in toxicology science and policy are inextricably linked.  For example, the FDA (2014) fish risk-benefit analysis contains several tables of ‘assumptions vs implications’ that were introduced because the Office of Management and Budget thought they were necessary to explicate potential biases.  But, sometimes the assumptions were essentially indisputable or had no clear political implications associated with them – so we tried to make some implications up.  In fact, the linkage between science and policy isn’t real at all – it’s done by design.  The biology underlying cancer and other diseases is very complicated, and the dose-response models typically used to generate estimates are approximations at best.  Trying to manipulate them to achieve a predetermined result tends to be obvious to anyone who isn’t doing the same.

But, make no mistake; the biology is complicated.  While the dose-response function used to characterize the risk from arsenic in rice serves as a nice exemplar of what cancer guidelines could be, the assessment itself is far from perfect.  There are many scientific issues yet to be addressed, some which are unique to arsenic.  For example:

  • The models used in the rice assessment are standard models that are in EPA Benchmark Dose modeling software.  Those are perfectly adequate for characterizing the range of what the risks associated with exposure to arsenic might plausibly be, but I would hesitate to say that they are up to the task of characterizing what the risk is most likely to be – no matter how the alternative models are weighted.  That problem might be addressed by adding one or more complex biological (i.e. toxicokinetic/toxicodynamic) models to the probability tree, and perhaps eliminating the models there now.  
  • The dose-response models in the apple and rice risk assessments are based on the analysis of the single cohort.  While the single studies chosen were the best available, a more concerted effort should develop a dose-response model that is reasonably consistent with all reported observations.
  • There is evidence that exposures earlier in life are more important.  The rice dose-response model dealt with that issue by just focusing on exposures under the age of 50, but a model that parameterizes the temporal component of the cause-effect relationship would be far superior (i.e. a time-to-tumor analysisof some sort).   At least theoretically, it should be possible to produce a model that is consistent with both the data from Taiwan (that is better for characterizing the influence of dose), and the results from Chile (that are more useful for characterizing the influence of age at time of exposure).  However, Individual subject data would be necessary to do that.

There are a host of other nagging scientific details that could be dealt with; but they only become important when the goal is to produce estimates that scientifically defensible, as opposed to simply conforming to a default procedure justified solely by agency policy.  The EPA has far more resources to devote to dose-response analyses than the FDA did.  If they can awaken from the dogmatic slumber induced by their own guidelines, perhaps they can do much better.

References

Carrington CD, Murray C, and Tao, S. (2013). A Quantitative Assessment of Inorganic Arsenic in Apple Juice.
U.S. Environmental Protection Agency (1986).  Guidelines for Carcinogen Risk Assessment.  EPA/630/R-00/004
U.S. Environmental Protection Agency (2005).  Guidelines for Carcinogen Risk Assessment.  EPA/630/P-03/001F.
U.S. Food and Drug Administration (2016).  Arsenic in Rice and Rice Products: Risk Assessment Report

Official Post Soundtrack


Rundgren, Todd (1973).  Just One Victory.  In: A Wizard, A True Star, Track 19.


Post Notes

Thesis Post #60.  It also the first of a two part commentary on the Arsenic in Rice and Rice Products: Risk Assessment Report released by the FDA last week.  In addition, it is an executive summary of many other blog topics, which is why there are many links to other essays.