Tuesday, March 31, 2026

To Bayes or Not to Bayes, That Is the Question


The Responsibility Gap

There is no reason to assign probabilities to competing theories unless there is an actual decision that must be made.  Pascal knew that, and that’s one of the things you would learn from chapter 6 (Hacking, 1975).  It obviously can’t be a flawless process; otherwise the correct theory would be assigned a probability of 1 every time.

So, I figure the next best thing is to have probability assignments that are scientifically defensible.  At least that’s what I tried to do when working for a regulatory agency as a scientific advisor because a) I figured it was my job, b) I liked the job, and c) I was allowed to do it.  I always thought that would start a conversation about what the probability assignments should be.  But that didn’t happen for several related reasons:

1) It may be easier to not make a decision at all.  It’s sort of an FDA tradition to declare emergency and then at a later date declare victory without doing anything at all.
2) If a decision absolutely must be made then it is much easier to with a formula, e.g. safety/uncertainty factors that doesn’t need to scientifically defensible.
3) If a risk estimate is necessary then it’s much easier to use a default assumption (linear extrapolation from high doses

All of those techniques insulate the expert from the decision, which obviates the need to assign probabilities to alternative hypotheses.  

Bayesian metholodogy also insulates experts from the decision, but to a lesser degree.  Since it does lay out the competing theories that underly a decision, I think it is far preferable to the backroom methodologies outlined above.  But perhaps the best about Bayesian methodology is that the practitioners actually want the job.

Help Wanted

I started assigning probabilities to alternative hypotheses because I figured it was my job and for many years I was allowed to do it.  I was also allow to participate in the writing of my own job description and I always made sure it said I was supposed to “convey uncertainty about potential health effects arising from contaminants  in food to decision makers”.  But that didn’t last; after a reorganization my job description changed to something like “support a decision that has already been made”.   That’s when I pretty much figured I’d rather write a blog than work at the USFDA.

I don’t really want my old job back.  I’m too old for that.  But I still figure someone else should have it.  There’s been another reorganization, but there is still a contaminants branch facing the same old problems.  Seems like a lot of people in the EPA and WHO should know their way around a probability tree as well.  

But it isn’t really just a government issue; it’s primarily a science problem.  The Bayesians shouldn’t need to identify plausible theories after a study has already been conducted.  That should have happened before the effort to design experiments and/or collect observational data began.  But of course, it is entirely possible that all of the hypotheses considered at the outset of a study are disproven by the new data.  That makes it time for a new theory.  Neither a frequentist or Bayesian analysis can help with that.  But a probability tree can.

Model Shopping

Perhaps the scientists who need probability trees the most are epidemiologists, especially the ones doing multivariate analysis with multiple putative causal influences.  I’ve been over some of it before from a historical perspective (Neyman was right, but Fisher sold more textbooks); null hypothesis testing doesn’t necessarily test the hypotheses that really matter.  That can easily set up a model shopping exercise that is solely interested in generating statistical significance, possibly by using a model that isn’t plausible in the first place.  I’ll also add the general point made in my last post that trying to turn hypothesis testing into a statistical exercise that treats observations as instances in stochastic probability theory rather than evidence for or against a theory even when the underlying theory isn’t stochastic at all. 

Statistical significance testing isn’t crazy when the number of alternative theories is exactly two, and the number of observations is small.  But otherwise, it’s nuts.  Model shopping isn’t such a bad idea when you are shopping for plausible theories.  However, it needs to happen as part of an open discussion that even someone working for the government can take part in.  That means the set recorded observations  used in published studies needs to be shared.  Furthermore, the search for more plausible theories doesn't stop just because a paper has already been published.  

Reference

Hacking, I (1975).  The Great Decision.  In: The Emergence of Probability.  Cambridge University Press, pp. 63-72.

Official Sound Track

Talking Heads (1977).  Don't Worry About the Government.  In: Talking Heads: 77, Track 8.


Tuesday, March 10, 2026

An Ode to Regression Analysis

Hello Again

It's been a while.  But the bug to write yet another essay has bitten me and I don't know what to do with it besides putting here.  It more or less started with having a wiki article on the History of Probability called to my attention.  I was gratified to see that it opened by acknowledging the duality of probability that I figure is a matter of psychological fact.  But then, as usual in my experience, the rest of the article proceeded to focus on stochastical probability, aka frequency of occurrence.  
 
Most annoying to me is that even though Ian Hacking’s The Emergence of Probability was referenced three times, the chapter (8) that discussed Pascal’s Wager wasn’t mentioned at all.  That’s where Hacking discussed the ins and outs of using probability trees (aka probability logic) to represent the evidential status of competing theories.

I get it; frequentist probability is much more amenable to a mathematical treatment than the evidential ancillary could ever hope to be; and that’s what the article is really about.  Hume (1739) and Hill (1965) resorted to subjective rules of evidence with no mathematics whatsoever rather than a deductive process where the premises automatically dictate the conclusions.  That doesn’t fit into a history of mathematical probability.

But I've been over all that before on this blog.  What's got me going again is that I've come up with another angle.  Quantifying stochastic probability has been done, but quantifying evidence is another thing altogether.

An Important Caution

Devising a mathematical methodology for assigning a numerical probability to competing theories is rather obviously not always possible.  For example, Pascal’s wager was on “God Is” vs “God Is Not”.  Deciding who the killer is in a murder mystery doesn't involve mathematics either even when it is quantified beyond a reasonable doubt.  Furthermore, mathematics is a deductive process; the conclusions must follow from the premises.  On the other hand, weighing evidence or generating hypotheses in the first place is inductive.  Or so they say, because it is also usually conceded that weighing evidence involves subjective judgment.  But when the theories themselves are mathematical then perhaps something useful can be done to make the evaluation not quite so subjective.  

Regression Analysis

As it turns out, if you ascribe discrepancies between a quantitative model and a set of observations to measurement error with a stochastic distribution then you can turn the estimation of model parameter values into a statistical problem.  It’s neat trick.  Sure, you can use a ruler to draw a straight line through scattered data, but different line drawers may end up with somewhat different slopes.  But with least squares linear regression you get the same result every time. 

Linear regression can performed with relatively straightforward mathematics.  But the model has to be linear.  However, with the aid of a computer using trial and error methodology where parameters are adjusted up and down to find if the fit improved or not, least squares regression can applied to any model.  Furthermore, it doesn’t even have to be least squares; any other methodology that weights the relative importance of discrepancies between model and observations may be used instead.  In particular, weighting unsquared residuals places less weight on large deviations than squared residuals do.

However, the underlying rationale for regression analysis is not entirely justified.  First, even if the discrepancies between model and observation are a result of measurement error the actual distribution is usually unknown.  Second, the set of observations may not be entirely representative of the actual distribution.  Third, the model may be not entirely correct.  That is especially likely with multivariate analyses where mismodeling one quantitative relationship can end up with misestimation of the other model parameters as well.  At that point it may be time to consider a new hypothesis.

So calling regression analysis “statistical” or even mathematical is a big stretch.  But I still think it’s very useful because it is using data as evidence for models and theories.  In fact, it can be thought of as quantitative induction.  That is good, very good in fact.  Furthermore, it seems clear that regression analysis has a role to play in filling out probability trees for competing theories with numerical probability assignments that sum to one.

Quantifying the Probability of Competing Hypotheses

The Bayesian Strategy

Employers of Bayes Theorem definitely understand that probability is not the same as frequency of occurrence and they are also comfortable with assigning probabilities to competing models.  Known as Bayesian inference, this is accomplished by assimilating the alternative models into a supermodel and then performing a regression analysis that assigns greater probabilities to the model(s) that fit the best.  

A Beyesian analysis can  also let subjective expert opinion be part of the process, but there’s a catch; the contributions from the experts comes before the regression analysis.  That suffers from the same general problem of trying to make grading evidence a deductive exercise; it’s just not consistent with the way science works.  After all, the issue underlying hypotheses lies in evaluating if they are true rather than how often they are true.   

The Pearson Strategy

There are two sorts of correlation coefficients, aka r-values.  The first measures the association between two different measurements of sets of observations  (e.g. genetics and the occurrence of a disease).  But a Pearson correlation coefficient can also be used to measure the relationship between the values predicted by a model and those observed, and it’s generated by linear regression.  You can easily produce something analogous to the Pearson r value for any regression methodology. The Bayes factor is also functionally equivalent to an r value.

Pearson himself thought the r value was useful for grading the strength of an inference (Porter, 1986).  It plugs in nicely to the first three Hill criteria, namely strength, consistency, and specificity.  It’s not to hard to argue that a model or theory with a higher r value deserves a higher probability assignment.  You could even devise an algorithm or equation that at least somewhat fairly directs the relationship.  Yes, it would be somewhat arbitrary, but I’ll take it all day over safety factors or default assumptions.

In Summary

There are two approaches for combining data and expert opinion.  The Bayesian approach starts with expert opinion and then uses data to produce final evidential judgments.  The Pearson-Hill approach produces a measure of how well the data fits each hypothesis, but leaves the final evidential judgment to experts.  I'll discuss pros and cons next.

References

Hacking, I (1975).  The Great Decision.  In: The Emergence of Probability.  Cambridge University Press, pp. 63-72.

Hill, AB (1965).  "The Environment and Disease: Association or Causation?". Proceedings of the Royal Society of Medicine. 58 (5): 295–300.

Hume, D (1739).  A Treatise on Human Nature.  Book I, Section XV.

Porter, TM (1986).  The Rise of Statistical Thinking 1820-1900. Princeton University Press.

Official Sound Track

Beatles (1969).  Come Together.  In: Abbey Road, Track 4.


Thursday, August 12, 2021

The Mantra

 I first heard it early in my career in my FDA career at an EPA symposium on manganese.  It went something like this:

“Where health is concerned, money doesn’t matter”

What an utterly stupid thing to say, I thought.  But to my consternation, many voices around the room followed with a “hear, hear”.  I have heard the mantra many times since, and I think it also lies unspoken behind many public policy decisions.  I have come to realize that it isn’t quite so stupid when uttered by people who are in the business of providing health benefits.  What they are really saying is:

“Where health is concerned, give us all your money”

Health Care

While I worked at the Center for Food Safety and Applied Nutrition, I always thought the mantra was uniquely associated with toxicological issues in food safety and environmental regulation.  But I’ve been out of the business for over six years now, and I now realize that the Mantra has a much wider presence.  In particular, the phrase “access to health care” sounds suspiciously like an alternative version of the Mantra.   What is “access” supposed to mean?  It obviously doesn’t mean everyone will be entitled to any and all medical procedures regardless of cost.  I think what it really means is:

“Give us all your money and we’ll decide what to do with it” 

But to my consternation, that seems to be exactly what we (in the US) are doing.  All the unfunded mandates (Obamacare, the hospital mandate, the employer mandate) all funnel money into the pockets of the medical industry with no consideration of how the money will be spent.  It’s why we spend twice as much money as any other country in the world and get less for it.  It’s socialized medicine run for profit; the worst of socialism and capitalism all rolled into one.   

However, Medicare is a different story.  Since the government must work with a limited budget, they are forced to be at least semi-rational about how the money is spent.  That’s why I think replacing the unfunded mandates with a fiscally conservative version (no we don’t need to spend more money on health care at public expense) of  Medicare for All is a mighty fine idea.  Makes no sense to have socialized medicine for poor people and old people while not giving it to the people who work and pay for it.   

COVID

But the poster child for the Mantra has to be COVID.  While I think Tony Fauci is a mighty fine scientist, that doesn’t mean he should be in charge of managing the economy.  I believe the initial reaction to the pandemic was overreaction driven by the Mantra.  Plus the aftermath reminds me of the fate of oyster beds after an oil spill; once the bureaucracy has taken control, it doesn’t want to let go until it has attained some arbitrary safety standard that never existed before.  

I do believe that people have generally gotten more rational about COVID, but there’s still a long way to go.  There are money, freedom and health tradeoffs with each and every mitigation technique.  We need to make the hard choices about which ones are really worth it.  I think vaccinations and masking are generally worth it, even if a mandate is required.  Since it became apparent early on that asymptomatic people can transmit the virus, haven’t thought contact tracing could work for a long time.  I don’t quarantining is especially worthwhile either.  I wonder if it’s wise to let hospitals fill up with COVID patients when they may have more pressing business to attend to.   Which brings us to public gatherings, especially schools.  

Looks like we are going to have to experiment a bit; let’s not hum the Mantra.

Official Post Soundtrack

Killing Joke (2005).  "Medicine Wheel."  In: Democracy, Track 8.


Tuesday, April 23, 2019

A Minority Opinion

Preamble

I haven't posted in over two years, and that's largely because I hadn't anything further to say, and that fact is attributable that I spend most of my time these days doing other things.  But, I still get dragged into it sometimes.  In particular, I've been participating in a World Health Organization workgroup for the last year and a half that was intended to come up with codifying standard practice for benchmark dose modeling.  That involved phone conferences and critiquing text written by other members.  Things weren't really progressing towards any sort of a conclusion, so a meeting was convened in Geneva last month that also brought in a number of other participants.

As I pretty knew already, among the initial workgroup members, I was in the minority in at least two respects.  First, while I'm a pharmacologist/toxicologist by training, the other members are largely statisticians.  Secondly, while I am very interested in dose-response modeling, I am actually not all that keen at picking a point on the curve to be the "Benchmark Dose".  Even though doing so might be useful sometimes, it never seemed to help with the problems I worked on at the FDA.

I wrote a short one page essay the morning after I got back. It was ostensibly written for inclusion somewhere in WHO document that is to be the end product of the meeting.  However, I don't know if or when that will happen.  So, I'll share it here just case it never goes anywhere else.

Dose-Response Modeling and Weight of the Evidence

Causal relationships can be expressed mathematically, and the expression of acceleration attributable to gravity is perhaps the most well-known example.   Dose-response models are quantitative expressions of causal relationships in pharmacology and toxicology.  However, even when it is expressed mathematically, the validity of the expression of causality ultimately depends upon a judgment that is not itself mathematical (Illari and Russo, 2015).  In the fields of medicine and physiology, perhaps the best of evidence of that comes from the fact that when Hill (1965) gave his widely known lecture on causality before a group of statisticians, he used no mathematical equations whatsoever.

Weight of the evidence approaches have been used for dose-response modeling both at JECFA (e.g. for lead, WHO 2000 and WHO 2011) and elsewhere (e.g. Morgan and Granger, 1980; Evans et al, 1994, and Carrington et al, 2011).  Using weight of the evidence to address dose-response model uncertainties is largely the same as when Bayesian methods are used.  There is still a need to identify a finite set of alternative models or hypotheses, and the models are still either directly fit to data or designed to be consistent with the empirical record.  Furthermore, both approaches utilize expert opinion, and at the end of the process probabilities are assigned to each alternative model so that they all add up to 1.  However, there are important differences.

  • First, Bayesian methodology uses expert opinion prior to curve-fitting, and then “updates” the probabilities initially assigned by the experts as part of the curve-fitting process to yield the final model probabilities.  On the other hand, a weight of the evidence approach does not assign model probabilities until after curve-fitting has taken place; experts may use information about how well each model describes the data, but also use other theoretical and experiential criteria as well.  Because the Bayesian approach alters expert option after it is expressed, it has the potential of yielding final model probabilities that contradict what experts believe.
  • Second, because it is amenable to automation, Bayesian methodology is far more reproducible than a methodology which depends solely on expert opinion.  Model probabilities assigned by experts may vary among experts or even a single expert over time.   That fact perhaps makes the Bayesian methodology preferable when a standardized approach is desirable and there is no strongly held expert opinion.
  • Third, because it is thought of as a mathematical exercise, calculating Bayesian probabilities requires the use of models for which log-likelihood functions can be calculated.  For more complex biological models, that may not be possible.  Under those circumstances, consulting expert opinion is really the only option.

Although they have been used for other purposes (e.g. Suter and Cormier, 2011) a formal process of the same ilk as the Hill criteria (Hill, 1965) is not typically implemented for quantitative dose-response modeling.  A process for weighing evidence could temper differences of opinion among experts regarding dose-response model form without eliminating expert opinion altogether. 

Assigning probabilities by committee would also help.   In place of the “associations” that concerned Hill, one or more numerical goodness fit measures could be used to argue for or against specific models.  The other Hill criterion that is directly relevant to dose-response modeling is the requirement for a “biological gradient”.  Quite simply, a dose-response model ought to look like what a dose-response relationship is supposed to look like.  That criterion could perhaps be subdivided into theoretical and experiential components.  As an instance of the former, an argument that a dose-response relationship cannot be supralinear as the dose approaches zero can be based on the notion that it violates the generally accepted biochemical law of mass action (Tallarida and Jacob, 1976).  An experiential argument would reflect the experience of toxicologists with other analogous dose-response relationships. 

References

Carrington CD, Murray C, and Tao, S. (2013). A Quantitative Assessment of Inorganic Arsenic in Apple Juice.  

Evans, J. S., Graham, J. D., Gray, G. M. and Sielken, R. L. (1994), A Distributional Approach to Characterizing Low‐Dose Cancer Risk. Risk Analysis, 14: 25-34.

Hill, Sir Arthur Bradford (1965).  The Environment and Disease: Association or Causation?  Proc Royal Soc Med 58:295-300.

Illari P and Russo F (2015).  Chapter 6: Evidence and Causality.  In: Causality: Philosophical Theory Meets Scientific Practice.  Oxford University Press, Oxford, pp. 46-59.

Morgan, M. G., Morris, S. C., Henrion, M. , Amaral, D. A. and Rish, W. R. (1984), Technical Uncertainty in Quantitative Policy Analysis — A Sulfur Air Pollution Example. Risk Analysis, 4: 201-216

Suter, GW and Cormier SM (2011).  Why and how to combine evidence in environmental assessments: Weighing evidence and building cases.  Science of the Total Environment 409:1406–1417.
Tallarida RJ and Jacob LS (1979).  Chapter 3: Kinetics of Drug-Receptor Interaction: Interpreting Dose-Response Data.  In: The Dose-Response Relation in Pharmacology.  Springer-Verlag, New York, pp. 49-84

WHO, 2000.  Lead, 53rd JECFA

WHO, 2011.  Lead, 73rd JECFA

Post Note

I do believe that Bayesian Model Averaging will work reasonably well for the evaluations where the risk is fairly trivial.  However, for more serious purposes (e.g. arsenic, lead, etc) it is a step in the right direction, at best.  But you just can't throw expert opinion, common sense, and associative learning into the dustbin of history by burying it in a prior.  It's silly, and sometimes obviously so.

Sunday, August 14, 2016

Individual Fish Risk Benefit Model

This page is set up as an adjunct to the discussion in The Science-Policy Shell Game concerning fish consumption advice.  I may replace the Excel macro linked here now with something prettier and/or easier to use at some point in the future.

Software

This Excel macro, Personal_Seafood_Net_Effect_Estimator.xlsintegrates four components presented earlier:




Sunday, July 10, 2016

SPSG #13: Ending the Game

This chapter outlines strategies for overcoming the difficulties noted in earlier chapters. First, adopting some common legal strategies for separating Matters of Fact from Matters of Law could do wonders. The creation of job positions for Science Judges whose sole responsibility is to disentangle science matters from policy matters could facilitate that. Second, while eliminating the science-policy shell game entirely is probably not possible, there is no reason why it should be condoned or institutionalized. Therefore, the EPA assessment guidelines for cancer and noncancer endpoints both need to be rewritten. As the face of the Safety Assessment Paradigm, the Reference Dose background document should be rewritten to make it clear that the product of the assessment is a regulatory policy rather than a statement of scientific fact. As for the cancer risk assessment guidelines, instead of enshrining the default option with the Point of Departure, the guidelines should use probability trees to make the default option going away entirely. Furthermore, there is no reason to the restrict the use of quantitative risk assessment to just cancer. However, solving that problem will create another problem: At least in public health, a legacy of the institutionalized use of the science-policy shell game has virtually eliminated risk management as a federal job position. So, doing a risk assessment is of very little use if no one in the federal government has the responsibility for managing issues. That can happen, but position descriptions will need to be rewritten. Finally, research should be funded to support science, instead of supporting technocratic shell games. In particular, the enterprise of environmental epidemiology needs to be redesigned. Statistical significance testing should be eliminated as the primary means of drawing conclusions from data, and studies should be designed to increase or reduce evidentiary weight accorded to causal theories instead. Perhaps most importantly, observational data needs to be shared. When it comes to analyzing data, regardless of what their source of funding is, investigators cannot be given complete deference in conveying what the data infer. As an academic recommendation, teaching Statistics and Probability as separate subjects would clear up more than a few nagging philosophical problems.  Finally, the facade of impersonal scientific objectivity needs to be abandoned. Scientists should be both free to speculate and humble enough to admit their theories may be wrong.

SPSG #12: Personal Technique

This chapter is largely written in the first person, and it does so largely for the purpose of disparaging the concept of scientific objectivity. It starts out by describing how several reorganizations dramatically affected the branch at the USFDA that I worked in for 25 years. In the end, it was swallowed up by the shell game. It then goes on to discuss the importance of recognizing the subjective nature of science, particularly when the science is unsettled and uncertain. The objectivity facade is partly attributable to scientific writing style that takes the author out of discussions of factual issues, which hides tha fact that personal opnions are beign expressed.   To demonstrate what science is really like, I walk through the personal choices I made in developing the dose-response model for arsenic and lung cancer that was used for the apple juice and rice risk assessments discussed in Chapter 11. Some of those choices were done by committee and some were not, but either way they all involved subjective scientific judgements made in a fog of uncertainty. In one case, I made a different choice that I had previously because new information influenced my subjective judgment about how to go about estimating lifetime risks from a prospective epidemiological study, which underscores the notion that “objective” reality evolves with scientific inquiry. The resulting dose-response model is then used to provide risk estimates for someone with a high-end (for the United States) arsenic intake. In addition to providing the lifetime risk estimates that were also given in the FDA reports on apple juice and rice, estimated changes in average life expectancy are also provided. For the purpose of making an individual choice, the latter measure is far more meaningful. The chapter then suggests that the inability of EPA to provide a dose-response characterization for arsenic may stem from a wrongheaded demand for objectivity that dictates the use of the wrong probability and the wrong personnel for a job that needs statistical theory instead of statistical probability.

Saturday, July 9, 2016

SPSG #11: The Technocracide

Since it kills the Safety Assessment Paradigm every time, it has always been very clear that arsenic would never make it as a food additive. Although arsenic commonly occurs in food as a contaminant, the concern for arsenic in food was always mitigated by the fact that the largest exposures have generally been from drinking water. That equation changed in 2001, when the EPA passed a regulation for arsenic in drinking water that changed that equation. Because the drinking water rule required a cost-benefit analysis, the decision process was supported by a risk assessment that produced risk estimates; in spite of the guidelines, it was consistent with the risk assessment paradigm. However, the Office of Water did find it necessary to hire outside consultants to accomplish that goal. In any case, arsenic exposure from water was reduced and arsenic in food became a relatively bigger issue as a result. As a result of public attention in 2011, the FDA issued guidance "action" levels for arsenic in apple juice in 2013 and rice in infant foods in 2016. From a risk management standpoint, both efforts were abject failures. Although risk assessments were produced, they didn't really support the guidance in either case. At least part of the reason is that setting levels usually isn't an effective way to manage risks from contaminants. Preventing something from getting into the food in the first place can be far easier, but that isn't always possible. Nonetheless, the FDA went ahead anyway with action levels anyway. In the case of apple juice, it isn't too hard to figure out why; the FDA commissioner publicly promised an action level before the risk assessment was done. The reasoning that went into the rice guidance is more mysterious. Even though there is nothing in the risk assessment to indicate that they are uniquely susceptible to arsenic, the FDA advised both infants and pregnant women to reduce their rice intake, but gave no advice for anyone else. In fact, the exposure assessment indicated that the greatest exposures to arsenic from rice are in adult males. On a more positive note, the FDA cancer risk assessments for both apple juice and rice solved the default option problem by using probability trees to represent the theoretical probability associated with the dose-response relationship for arsenic and both lung and bladder cancer. The main lesson to be learned from those exercises is that regardless of how well an assessment represents current science, if the message takes precedence over the result, there will be no reason to expect public health to improve.

SPSG #10: The Paradigm War

Methylmercury in fish has been a major issue for both the EPA and the FDA since several epidemics occurred in Japan and Iraq in the 60s and 70s. At about the same time (early 90s) as the FDA started quantifying risks for methylmercury in fish and issuing consumer advice for commercial seafood, the EPA started giving recreational fish consumption advice based on the EPA Reference Dose (RfD). In 1999 congress asked the National Academy of Science (NAS) to evaluate the RfD for methylmercury, and a report was issued in 2000. The fact that congress even asked the question of the NAS cemented in many people minds that the RfD was and is a statement of a scientific fact. That meant that if it was true for EPA then it had to be true for FDA too, and as a result any attempt to quantify the risks and provide information about what the risks came to be viewed as a political attempt to undercut the science. By 2004, the FDA and EPA had agreed to give joint advice for fish consumption, but there was no agreement about what the basis or the rationale for the fish advisory was. While the EPA thought the RfD was paramount, the FDA chose to pursue a quantitative strategy that balanced benefits and risks; so the fish-risk benefit assessment was an FDA-only affair. But perhaps the most important difference was about what information, if any, would be given to consumers. The RfD treats consumers in the same manner as it treats agency managers; it decides for them, and as a result there is no basis for providing the information to consumers that will let them decide for themselves. As an alternative, a few representative risk estimates are provided for the consumption of fish during pregnancy using a version of the risk assessment model developed for the FDA that is designed to estimate risks for individual consumers.

Tuesday, July 5, 2016

SPSG #9: A Practical Guide to Theoretical Probability

This Risk Analysis methodology chapter is the applied version of the philosophical discussion of probability presented in Chapter 2. It also fixes the flaws in the Redbook paradigm discussed in Chapter 4, resulting in the Guillotine paradigm. It begins with a discussion of characterizing uncertainty when there are both statistical and theoretical probabilities involved. While a theoretical probability does not need to be quantified when it is the only probability involved or when there is no decision at stake, giving it the same epistemic standing as a statistical one is unavoidable when both matter. However, that does not mean a theoretical probability can be used as if it were a statistical probability. A theoretical probability is perhaps true always or perhaps false always; is it not true sometimes and false at other times. The discussion then turns to the problem of assigning probabilities to alternative theories. Declaring that all sum to one is a simple matter, but deciding the probability of each theory is not. Since theoretical probabilities are subjective, depending on the opinions of those who have one (that usually means experts) is in some way is inevitable. However, instead of asking experts to assign probabilities to theories directly, there are advantage to garnering opinion in the form of evidential weights, where each alternative theory is evaluated more or less independently. Although formal weight-of-the-evidence schemes have been developed for many regulatory purposes, they are not usually thought of as quantitative exercises. However, it has been done and could be done better. It is also argued that WoE analysis and dose-response modeling need to be more tightly integrated, especially when the judgment that there is a causal relationship becomes more likely than not. First, the shape of the dose-response relationship may influence the judgment that there is a causal relationship. Second, the last vestiges of causal uncertainty may not matter if the estimated risks are too low to matter or high enough to be a concern even if they are only probable.

Monday, July 4, 2016

SPSG #8: The Wrong Probability

This chapter is about the problems associated with using statistical probability as the only probability, especially in epidemiology. While different scientific disciplines typically rely on somewhat different collections of convincing arguments, the conduct of epidemiology can aptly be compared to a trial for murder. Since the issue is causality, theoretical probability is front and center. Yet, at least when designing studies and publishing studies, environmental epidemiology studies often rely on tests of statistical significance testing for drawing conclusions. Arthur Bradford-Hill disparaged this practice over 50 years ago, and he is still quite right; statisticians are using the wrong probability. But that isn't the only problem. Epidemiologists (or their statisticians) often treat measures designed to quantify strength of association for the purpose of arguing causality as if they were measures of effect; thereby completely missing the point of having them in the first place. Next, epidemiologists are often reluctant to share raw data. While there are many possible explanations for this practice, the fact that other analysts would be able to use the data to explore and support theories not utilized in the published report is chief among them. The data sharing problem becomes especially evident when the theories used in published analyses are obviously wrong, either when they are first published, or perhaps later. This more or less forces the court of scientific opinion to rely on hearsay evidence. As an example of that, the use of log transformed measures of dose in multivariate regression analyses is discussed. Since it is an established analytical procedure, there is a tendency to think of regression analysis as a "theory-free" analysis that provides conclusions that are largely empirical. But, that isn't true at all. Linear regression analysis presumes that the quantitative dose-response relationship is linear. Similarly, doing a linear regression analysis with the log of dose presumes that the quantitative dose-response relationship is loglinear. But that results in a supralinear function where not only do the effects get bigger as the dose gets smaller, the effect approaches infinity as the dose approaches zero. Even though that's quite impossible, the practice continues, and that is probably because testing a theory that is definitely wrong is a reliable way of producing scary statistically significant low dose effects.

SPSG #7: The Sociology of Technocracy

This chapter is like a sociology of science essay, except that it is really about politicians in lab coats, with most of it being concerned with toxicology. Many of the roots of the SPSG can be found in academia. First, there is a discussion of the Information Quality Act of 2002 that sought to make the information used by the federal government more objective. Yet the Office of Management and Budget interpreted "objectivity" as meaning "peer reviewed". While that could potentially subject scientific claims to cross-examination by outside experts, without separation of science and policy, peer review can also be used to prevent cross-examination altogether. While toxicology initially was primarily associated with the drug industry, it has become increasingly concerned with environmental regulation. The growth of environmental toxicology programs that are almost entirely concerned with government as a career path are a prime example. As a result, environmental toxicologists can potentially complete their careers by only talking among themselves. Although the Society of Toxicology did not have a Code of Ethics when it was formed in 1961, it does now. Many of the recommendations seem to be political statements that on closer examination aren't necessarily ethical at all. In particular, members are required to be "advocates of public health" and "Abstain from professional judgments influenced by undisclosed conflict of interest". Both of these statements favor the interests of public sector members (i.e. technocrats) over those with private interests. Nutritionists can be technocrats too, and when nutrients are also toxic, that can create a clash of technocratic cultures. While nutritionists don't use safety factors, when it comes to considering dose-response relationships, their tradition is quite limited, perhaps by design. In the “Risk Analysis Paradogm”, Risk Communication is often recognized as a third component of the regulatory decision process, along with Risk Assessment and Risk Managment.  However, the roots of the discipline lie in the study of consumers responses, which makes it well suited for selling a decision that has already been made. Quantification can be part of the SPSG too.  Because of their ability to seemingly automate a decision process, often by ignoring or assiduously hiding theoretical probability, statisticians can play the shell game too.  Since they often equate the behavior of scientists with science, sociologists sometimes seem to backhandedly endorse the shell game.

Note

Although it wasn't my original intent, this chapter does read like a populist manifesto.  

Sunday, July 3, 2016

SPSG #6: Two Charades

While the earlier chapters are largely historical, this is the chapter that begins to speak of current practice, and it is also gives the book its title. Quite simply, the Science-Policy Shell-Game (SPSG) is a technocratic game played by treating science and policy as if they were interchangeable in both directions. It is comprised of two components. On the one hand, a statement is purported to be science in front of a political audience. On the other hand, the same statement is purported to be policy in front of a scientific audience. As the end result, a regulatory decision is shielded from both scientific and political scrutiny. Although there were many examples of it before, the SPSG was deliberately institutionalized by a committee of EPA technocrats in 1986. The object of their creation was the EPA Reference Dose (RfD). As an example of the Safety Assessment Paradigm, the RfD was no different from the ADI, except for one thing: It was claimed to be a scientific fact. As a further technocratic assault, the 2005 EPA Cancer Assessment Guidelines replaced plausible worst-case estimates with what amounted to a ban on theoretical reasoning. That was accomplished by interposing a "Point of Departure" between toxicological theory and regulatory decision making. That made no sense then, and it still doesn't. Why the SPSG is played is debatable; it may be some combination of agency managers hiding decisions they would rather not defend, scientists who don't want to cede regulatory control to agency managers or elected officials, statistical decision theorists who claim to make decision processes objective, or it may just be career maintenance and research dollars. But whatever the reasons are, the SPSG is antiscientific and antidemocratic. It cuts off scientific discussion from policy making with technocratic short cuts, and it leads scientists and the public to believe that the decisions faced by the government and themselves are far simpler than they really are.  The “Fifth Branch” is introduced as a term to describe the technocrats inside and outside of the federal government who play the SPSG.

SPSG #5: Dose-Response Theory

This chapter is a compendium of pharmacology and toxicology theory, and since it doesn’t build on any of the previous chapters, it is essentially a third introductory chapter.  Although does get a bit technical, it is designed to give a sense of what the quantitative issues are without delving into mathematics.  Although it isn’t necessary for most of the later chapters, it is provided as background material for some of the discussions involving theoretical probability in some of the later chapters, especially eight through eleven.  The chapter commences with a survey of basic concepts including biochemical mechanisms underlying the interaction of a toxic chemical with a biological molecule, and toxicokinetic theory that describes what happens before it gets there.  There is also discussion of statistical theories like probit analysis that treat the causal issue as a problem of describing how much the dose required to produce a given effect varies in a population.  There are also hybrid or two-dimensional models that describe the both magnitude of individual effect and population variation as well.  It has long been recognized that there is a temporal component to dose response relationships that can vary between both chemicals and effects; yet the temporal component is often ignored.  All said, biology is complex and dose-response theory is imperfect; there is plenty of room for improvement.  Nonetheless, one thing seems clear; the dose makes the poison.  Effects tend to get bigger as the dose gets bigger, but not necessarily in proportion.


Saturday, July 2, 2016

SPSG #4: The Risk Assessment Paradigm

This chapter is about the Risk Assessment Paradigm (RAP) that was originally introduced to chemical safety in food and the environment as a mechanism for dealing with carcinogens in food. The advantages of the RAP were touted by a 1983 report from the National Academy of Sciences (NAS) that is often referred to as the Redbook. In particular, the RAP was widely heralded as a democratic alternative to the technocratic SAP. This was accomplished by distinguishing a risk assessment process that characterizes what the risks are and a risk management process that uses the information from the risk assessment to make a decision. However, the Redbook version of the RAP had some significant flaws that kept it from really accomplishing the goal of separating science from policy.  Therefore, this book also uses a more basic definition of RAP called "the Guillotine Paradigm", that is predicated on the separation of "Is" from "Ought": The scientific discussion of what is known is a separate endeavor from the policy decision of what ought to be done.  The problems with the Redbook were identified in subsequent NRC reports.  First, the Redbook paradigm begins with a Hazard identification step.  Since this is to be done by scientific experts, it gave them control over what questions are to be answered.  That problem can be addressed by viewing the RAP as an iterative process, where the policy deliberation identifies the questions, and the risk assessment provides the answers.  Veiwing the RAP as a dialogue istead of a monologue opens the decision process up to democratic participation.  Secondly, the Redbook embraced the notion of the “default option”, where theoretical probabilities were to be resolved by giving one theoretical alternative preference as a matter of regulatory policy.  That problem can be solved with a probability tree that acknowledges the theoretical alternatives.  Thirdly, the Redbook paradigm was derived from procedures used at the FDA for dealing with cancer and the Delaney Clause, and it was therefore sometimes interpreted as being applicable only to cancer.  Although the RAP is named after the analytical component, the fact that policymaking (aka Risk Management) is recognized as a distinct process is at least as important.  The RAP is especially indispensable when regulatory decision making requires a rational process where the risk is balanced against something else, such as a benefit, another risk, or the cost of avoiding it.

SPSG #3: The Safety Assessment Paradigm

This chapter is about the procedure originally developed for the premarket approval of pesticides and food additives, which is referred to throughout the book as the Safety Assessment Paradigm.  The SAP gave birth to the Acceptable Daily Intake (ADI), which was the daily exposure to a chemical that would be considered acceptable by the FDA. The ADI was originally calculated by dividing the highest dose found to not have an observable (statistically significant) effect in a laboratory study by a safety factor of 100. There are three key features of Safety Assessment that are especially worth noting. First, since the SAP delegates the regulatory decision to experts, it is thoroughly technocratic.  Second, premarket approval and the SAP were designed to be precautionary; a chemical couldn't be used until it was shown to be safe. Third, it is presumed that the way to limit exposure to a chemical is the set a level that the government will consider to be acceptable. Although the SAP has evolved from its 1954 introduction, it still retains its unmistakable premarket approval origins.


Friday, July 1, 2016

SPSG #2: Two Probabilities and Frequency Too

This is a history of philosophy of science presentation, and it sets up the terminology used throughout the rest of the book: If the first chapter is about "safety", this one is about "probability". There are two very different concepts of probability, both of which are very old and generally familiar. From an etymological standpoint, the legal concept of probability dates back to Roman law. In common parlance, the evidential form of probability is being used whenever a proposition is said to be "probably true". Even though it wasn't called probability until Pascal and friends gave it that name in the 17th century, the concept of chance and its relationship to frequency of occurrence has been around since Aristotle. In common parlance, the probability of chance is being used whenever it is said that an event will "probably happen". In scientific and technical literature, it is the second "statistical" meaning of probability that is used almost exclusively, and that is almost certainly because it is more objective in an empirical sense. However, the other subjective form of probability, which is called "theoretical" or "evidential" probability in the book to emphasize its role in science, can still be found in scientific discussion all the time, but it usually appears in the words rather than the numbers.

There are two other important concepts introduced in this chapter as well. First, as a result of the statistical definition of probability, uncertainty and frequency of occurrence are often treated as identical concepts. That is incorrect both from a grammatical standpoint, and because even without the appendage of "probability", statistical theories concerning frequency of occurrence are important in many scientific disciplines. For example, in public health, regulatory issues often revolve around evidential probabilities of statistical theories. Second, abstract representations of probability are also sometimes referred to as probability themselves without reference to usage, which leads to yet another opportunity for semantic confusion. As they are taught in introductory courses concerned with probability and statistics, mathematical probability distributions that can be used to represent either frequency of occurrence or the probability of chance are well known. However, even though theoretical probability is clearly a matter of degree, it is much harder to quantify. A probability tree can be used to represent theoretical probability and provide a quantitative interpretation as well. The basic concept is very simple: The probability of all alternative theories or hypotheses under consideration sums to one. Scientific evidence may then be weighed in order to determine which hypotheses, and to what degree, are more probable than the others. Causal relationships are the most common issue in which scientific issues involving theoretical probability arise. For example, whether or not a particular chemical causes cancer, and if so how, involves theoretical probability. Although they have evolved somewhat, guidelines for establishing causal relationships in science have been around for centuries. They have been around in law for even longer than that.

Notes

Perhaps this should be the first chapter, but in the interest of not straining credulity at the outset I have put it second.

After having been at the FDA for several years, I ran into my first model uncertainty problem.  I traipsed over to the Office of Mathematics to inquire about the proper methodology for calculating the probability of a model being true. I asked a statistician who had been at the FDA for thirty years, and the answer I got astounded me:
You are not allowed to ask that question 
I obtained a second opinion from a younger statistician,  just trying to find out about how to go about identifying the best model among several and got an answer that I found to be no more satisfactory:
Find a biologist and beat it out of them
As a biologist who did not want to be beaten, I soon embarked upon a philosophy of science reading binge that lasted several years in the mid 90's.  The main thing I learned is that there is another probability that is quite different from the one the statisticians were using.  However, in the end I was unable learn very much about it from the philosophers.  This chapter is a compilation of the few gems that I managed to gather from my survey.  Perhaps not surprisingly, some of the most important insights came from practicing scientists and risk analysts. In any case, I managed to answer my own question, at least to my own satisfaction. Although beating the probabilities out of the biologists can be a fair characterization of the method, the beating can be avoided with a willing confession. The description of that solution begins in this chapter and finishes in chapter 9.

However, I ran into other difficulties subsequently.  It seems that not everyone wanted the problem to be solved at all.  The rest of the book is about that.


SPSG #1: Food Law and Chemical Safety

This is an historical review of food law concerned with chemical safety. It commences with a discussion of the 1906 Pure Food and Drug Act and ends with the Dietary Supplement amendments of 1994.  All told, there are about a half dozen statues that pertain to the regulation of chemicals in food by the U.S. Food and Drug Administration (FDA) that differ in many ways.  While virtually all of the laws governing the safety of chemicals in food require scientific interpretation, the way in which scientific expertise is to utilized to create a legal definition of what is considered safe or not by the agency necessarily differs among different statutes.  The statutes differ in the definition of harm, the burden of proof, and in their evidentiary standards. For the purposes of the rest of the book, the most important distinction is between additives and contaminants. Food additives are deliberately added to food, have an intended use, must be approved by the FDA before they can be used, and the burden of proof lies on the manufacturer. Because the approval process is structured and planned, the way in which scientific expertise is utilized also occurs in a somewhat predetermined manner. The evidentiary standard for arguing that an additive has not been shown to be safe is also very low; it must only be argued that there is substantial evidence that harm is possible. Contaminants are, by definition, present in food unintentionally and many occur naturally. Therefore, they don't have an intended use, they don't need to be approved, and the burden of proof is on the government to show that the contaminants is harmful enough to be worthy of regulation, and the judicial branch often makes the final determination of what will be considered "safe".  Other classes of chemicals fall soemwhere between those two extremes.  In particular, food additives in use prior to the Federal Food Drug and Cosmetic Act amendments of 1958 were exempted from the approval process.  This created a class of chemicals that are “Generally Recognized as Safe” where FDA approval is obtained by demonstrating history of use rather than going through a rigorous testing regimen.  As the end result, there is no consistent definition in either law or science about what the word "safe" really means.


Thursday, June 30, 2016

The Book

The Science-Policy Shell Game


Most of the previous discussion on this blog has been assembled into a ebook that is better organized and written than my more casual blog essays.  I have sprinkled a few new ideas into it as well.

Unfortunately, it isn't free.  Maybe that's because I think I can trick people into thinking it is worth reading by making them pay for it.  However, you can read the preface and summary for free, which should enable you to make and informed choice about whether or not it is worth five bucks to read the rest of it:

US Link, Search the title on Amazon elsewhere.

Chapter Links

I keep hoping that this blog will turn out to be a place where I can discuss the issues near and dear to my heart.  Towards that end, the following links are provided for comments and discussion of individual chapters:


Wednesday, May 11, 2016

Biological Problems

The Perils of Multivariate Linear Regression

As is often the case in the epidemiological literature on environmental influences on neurobehavioral development, Bowers and Beck (2006) noted that a paper by Lanphear et al (2005) “has suggested the existence of a supra-linear dose–response relationship between environmental measures such as blood lead concentrations and IQ”.  They then produced an analysis that indicates that the apparent supralinearity is an artifact resulting from the way the data were analyzed.  They stated their conclusion as follows:
Results of the analyses show that a supra-linear slope is a required outcome of correlations between data distributions where one is lognormally distributed and the other is normally distributed. 
While their mathematical analysis was indubitably correct, the way Bowers and Beck reported the results left something to be desired.  How the data are distributed is not really the issue at all.  Instead, the mathematical artifact they found results from conducting linear regression analyses with log transformed data.  If data from a normal distribution, or any other distribution, were log transformed prior to the regression analysis, then the same result would be obtained.  Furthermore, as demonstrated by Jusko et al (2006), a linear regression without log transformation with data drawn from a lognormal distribution does not result in a supralinear dose-response relationship.

To their credit, Hortung et al (2006) also understood that the real issue is the shape of the dose response relationship rather than the distributions that either the dependent or independent variables follow.  They therefore protested that Lanphear et al (2005) had considered the likely shape of the curve before conducting the regression analysis:
The shape of the exposure–response relationship was determined to be nonlinear insofar as the quadratic and cubic terms for concurrent blood lead were statistically significant (p < 0.001 and p = and 0.003, respectively).  Because the restrictive cubic spline indicated that a log-linear model provided a good fit to the data, we used the log of concurrent blood lead in all subsequent analyses of the pooled data.
But, there are many problems with this justification.  First, it is not at all clear how a spline analysis specifically supports a log-linear model, as opposed to other potential nonlinear models (e.g. a Hill function).  Second, there was no consideration of biological plausibility.  Like Bowers and Beck (2006), Lanphear et al (2005) seem to think establishing a causal relationship is a mathematical problem rather than a biological one.  Third, they did not consider the possibility that other covariates might explain the apparent nonlinearity.   Yet, off they went, and a dose-response model that predicts infinite large effects as the dose approaches zero was the inevitable result.  For all practical purposes, Bowers and Beck (2006) were entirely correct.  

Besides the fact that a loglinear dose-response model is a very poor theory, there is a more general lesson to be learned:  A multivariate regression with assumed quantitative relationships between the variables being modeled is highly prone to error.  While a loglinear relationship is obviously wrong, a linear relationship isn’t necessarily right either.  Correlations between variables may result in attribution of mismodeled causal effects to a variable that has no causal effect at all.  For example, if the relationship between socioeconomic status (i.e. the HOME score) and IQ is nonlinear with bigger impacts with low scores and negatively correlated with exposure to an environmental chemical, the some of the socioeconomic effect will erroneously appear to be a low dose effect attributable to the environmental chemical.  There are many other possible explanations as well, all of whicih are more probable than a dose response model than predicts incremental effects to get bigger as the dose gets smaller.

Process vs. Theory

Biological complexity often makes the pronunciation of definitive truths doe Medicine and Public Health practically impossible.  While relying on expert opinion is a common solution to that problem, that solution does not work well when opinion is divided.  As a means of coping with that problem, institutional decision making processes often employ structured evaluation systems to sort through what can often be a voluminous set of scientific literature.  The Safety Assessment methodology that is typically used for premarket approval evaluations is an example.   This description of Evidence-Based Medicine conveys the general ethos of such efforts:
Whether applied to medical education, decisions about individuals, guidelines and policies applied to populations, or administration of health services in general, evidence-based medicine advocates that to the greatest extent possible, decisions and policies should be based on evidence, not just the beliefs of practitioners, experts, or administrators. It thus tries to assure that a clinician's opinion, which may be limited by knowledge gaps or biases, is supplemented with all available knowledge from the scientific literature so that best practice can be determined and applied. It promotes the use of formal, explicit methods to analyze evidence and makes it available to decision makers.
There are two key concepts at work here.  First, the “beliefs of practitioners, experts, or administrators” are getting kicked to the curb in favor of “evidence”.  If you thought the beliefs of experts were based on scientific evidence, then you were misinformed, apparently.  Secondly, there is an emphasis on the “use of formal, explicit methods”, which also serve to limit subjective influences on the evaluation process. 

Experts are not always trustworthy, so the desire for a transparent process is entirely understandable.  But, getting a trustworthy process to replace the experts is easier said than done.  The process has to be designed by somebody, and that usually means experts.  There is also apt to be a negotiation process involved in getting the process to be accepted, so subjectivity isn’t really completely avoided.   But perhaps the bigger problem is that trying to deal with complex biological issues with a formula may often be rather stupid.  If all the studies show the same result, then it really isn’t going to matter whether the decision making process is expert-based or evidence-based.  If the results are different, then the systematic review may succeed at identifying the higher quality studies and grading the general result.  But, it won’t explain why the results are different.  It won’t figure out why a treatment may works sometimes, but not others.  That will take biological theories, and like the biases the evidence-based systems strive to avoid, those are subjective.   There are likely to be different theories, of course, and then the experts will inevitably get into a debate over which are more likely.  But, guess what, that’s the way science works: Trying to eliminate all potential bias with formulae will also eliminate scientific progress.

By all means, more transparency is needed.  In particular, let’s not trust authors to have the last word on how the data they have collected are analyzed and published.  Medical researchers and epidemiologists are notorious for not sharing original data involving human subjects, even when they are legally required to do so (Panhuis etal, 2014; Longo and Drazen, 2016).  That will allow better theories to flourish, and poor theories to flounder.

References

Bowers TS and Beck BD (2006).  What is the meaning of non-linear dose-response relationships between blood lead concentrations and IQ?  Neurotoxicology 27:520-4.

Hornung R, Lanphear B, Dietrich K. (2006).  Response to: “What is the meaning of non-linear dose–response relationships between blood lead concentrations and IQ?”.  Neurotoxicology 27:635

Jusko TA, Lockhart DW, Sampson PD, Henderson CR Jr., and Canfield RL (2006).  Response to: “What is the meaning of non-linear dose–response relationships between blood lead concentrations and IQ?”.  Neurotoxicology 27:1123–1125.

Lanphear BP, Hornung R, Khoury J, Yolton K, Baghurst P, Bellinger DC, Canfield RL , Dietrich KN, Bornschein R, Greene T, Rothenberg SJ,8, Needleman HL, Schnaas L, Wasserman G, Graziano J,13 and  Roberts R. (2005).   Low-Level Environmental Lead Exposure and Children’s Intellectual Function: An International Pooled Analysis.  Environ Health Perspect. 113: 894–899.

Longo DL and Drazen JM (2016).  Data Sharing.  N Engl J Med 374:276-277.

Panhuis WG van, Paul P, Emerson C, Grefenstette J, Wilder R, Herbst AJ, Heymann D, and Burke DS (2014).  A systematic review of barriers to data sharing in public health.  BMC Public Health 14:1144.

Official Post Soundtrack


Jackson, J (1980).  Biology.  In: Beat Crazy, Track 9.

Post Notes

Thesis Post #65.  This covers some of the same ground as Toxicology Meets Epidemiology, but with a more philosophical overview.