Feeling the Research

Daryl Bem must be sick of those puns by now.

Back in 2011 he published Feeling the Future, a paper that combined multiple experiments on human precognition to argue it was a thing. Naturally this led to a flurry of replications, many of which riffed on his original title. I got interested via a series of blog posts I wrote that, rather surprisingly, used what he published to conclude precognition doesn’t exist.

I haven’t been Bem’s only critic, and one that’s a lot higher profile than I has extensively engaged with him both publicly and privately. In the process, they published Bem’s raw data. For months, I’ve wanted to revisit that series with this new bit of data, but I’m realising as I type this that it shouldn’t live in that Bayes 20x series. I don’t need to introduce any new statistical tools to do this analysis, for starters; all the new content here relates to the dataset itself. To make understanding that easier, I’ve taken the original Excel files and tossed them into a Google spreadsheet. I’ve re-organized the sheets in order of when the experiment was done, added some new columns for numeric analysis, and popped a few annotations in.

Odd Data

The first thing I noticed was that the experiments were not presented in the order they were actually conducted. It looks like he re-organized the studies to make a better narrative for the paper, implying he had a grand plan when in fact he was switching between experimental designs. This doesn’t affect the science, though, and while never stating the exact order Bem hints at this reordering on pages three and nine of Feeling the Future.

What may affect the science are the odd timings present within many of the datasets. As Dr. R pointed out in an earlier link, Bem combined two 50-sample studies together for the fifth experiment in his paper, and three studies of 91, 19, and 40 students for the sixth. Pasting together studies like that is a problem within frequentist statistics, due to the “stopping problem.” Stopping early is bad, because random fluctuations may blow the p-value across the “statistically significant” line when additional data would have revealed a non-significant result; but stopping too late is also bad, because p-values tend to exaggerate the evidence against the null hypothesis and the problem gets worse the more data you add.

But when pouring over the datasets, I noticed additional gaps and oddities that Dr. R missed. Each dataset has a timestamp for when subjects took the test, presumably generated by the hardware or software. These subjects were undergrad students at a college, and grad students likely administered some or all the tests. So we’d expect subject timestamps to be largely Monday to Friday affairs in a continuous block. Since these are machine generated or copy-pasted from machine-generated logs, we should see a monotonous increase.

Yet that 91 study which makes up part of the sixth study has a three-month gap after subject #50. Presumably the summer break prevented Bem from finding subjects, but what sort of study runs for a month, stops for three, then carries on for one more? On the other hand, that logic rules out all forms of replication. If the experimental parameters and procedure did not change over that time-span, either by the researcher’s hand or due to external events, there’s no reason to think the later subjects differ from the former.

Look more carefully and you see that up until subject #49 there were several subjects per day, followed by a near two-week pause until subject #50 arrived. It looks an awful like Bem was aiming for fifty subjects during that time, was content when he reached fourty-nine, then luck and/or a desire for even numbers made him add number fifty. If Bem was really aiming for at least 100 subjects, as he claimed in a footnote on page three of his paper, he could have easily added more than fifty, paused the study, and resumed in the fall semester. Most likely, he was aiming for a study of fifty subjects back then, suggesting the remaining forty-one were originally the start of a second study before later being merged.

Experiment 1, 2, 4, and 7 also show odd timestamps. Many of these can be explained by Spring Break or Thanksgiving holidays, but many also stop at round numbers. There’s also instances where some timestamps occur out-of-order or the sequence number reverses itself. This is pretty strong evidence of human tampering, though “tampering” isn’t the synonymous with “fraud;” any sufficiently large study will have mistakes, and any attempt to correct those mistakes will look like fraud. That still creates uncertainty in a dataset and necessarily lowers our trust in it.

I’ve also added stats for the individual runs, and some of them paint an interesting tale. Take experiment 2, for instance. As of the pause after subject #20, the success rate was 52.36%, but between subject #20 and #100 it was instead 51.04%. The remaining 50 subjects had a success rate of 52.39%, bringing the total rate up to 51.67%. Why did I place a division between those first hundred and last fifty? There’s no time-stamp gap there, and no sign of a parameter shift. Nonetheless, if we look at page five and six of the paper, we find:

For the first 100 sessions, the flashed positive and negative pictures were independently selected and sequenced randomly. For the subsequent 50 sessions, the negative pictures were put into a fixed sequence, ranging from those that had been successfully avoided most frequently during the first 100 sessions to those that had been avoided least frequently. If the participant selected the target, the positive picture was flashed subliminally as before, but the unexposed negative picture was retained for the next trial; if the participant selected the nontarget, the negative picture was flashed and the next positive and negative pictures in the queue were used for the next trial. In other words, no picture was exposed more than once, but a successfully avoided negative picture was retained over trials until it was eventually invoked by the participant and exposed subliminally. The working hypothesis behind this variation in the study was that the psi effect might be stronger if the most successfully avoided negative stimuli were used repeatedly until they were eventually invoked.

So precisely when Bem hit a round number and found the signal strength was getting weaker, he tweaked the parameters of the experiment? That’s sketchy, especially if he peeked at the data during the pause at subject #20. If he didn’t, the parameter tweak is easier to justify, as he’d already hit his goal of 100 subjects and had time left in the semester to experiment. Combining both experimental runs would still be a no-no, though.

Uncontrolled Controls

Bem’s inconsistent use of controls was present in the paper, but it’s a lot more obvious in the dataset. In experiments 2, 3, 4, and 7 there is no control group at all. That is dangerous. If you run a control group through a protocol nearly identical to that of the experimental group, and you don’t get a null result, you’ve got good evidence that the procedure is flawed. If you don’t run a control group, you’d better be damn sure your experimental procedure has been proven reliable in prior studies, and that you’re following the procedure close enough to prevent bias.

Bem doesn’t hit that for experiments 2 and 7; the latter isn’t the replication of a prior study he’s carried out, and while the former is a replication of experiment 1 the earlier study was carried out two years before and appears to have been two separate sample runs pasted together, each with different parameters. In experiments 3 and 4, Bem’s comparing something he knows will have an effect (forward priming) with something he hopes will have an effect (retroactive priming). There’s no explicit comparison of the known-effect’s size to that found in other studies, Bem’s write-up appears to settle for showing statistical significance. Merely showing there is an effect does not demonstrate that effect is of the same magnitude as expected.

Conversely, experiments 5 and 6 have a very large number of controls, relative to the experimental conditions. This is wasteful, certainly, but it could also throw off the analysis: since the confidence interval narrows as more samples are taken, we can tighten one side up by throwing more datapoints in and taking advantage of the p-value’s weakness.

Experiment 6 might show this in action. For the first fifty subjects, the control group was further from the null value than the negative image group, but not as extreme as the erotic image one. Three months later, the next fourty-one subjects are further from the null value than both the experimental groups, but this time in the opposite direction! Here, Bem drops the size of the experimental groups and increases the size of the control group; for the next nineteen subjects, the control group is again more extreme than the negative image group and again less extreme than the erotic group, plus the polarity has flipped again. For the last fourty subjects, Bem increased the sizes of all groups by 25%, but the control is again more extreme and the polarity has flipped yet once more. Nonetheless, adding all four runs together allows all that flopping to cancel out, and Bem to honestly write “On the neutral control trials, participants scored at chance level: 49.3%, t(149) = -0.66, p = .51, two-tailed.” This looks a lot like tweaking parameters on-the-fly to get a desired outcome.

It also shows there’s substantial noise in Bem’s instruments. What’s the odds that the negative image group success rate would show less variance than the control group, despite having anywhere from a third to a sixth of the sample size? How can their success rate show less variance than the erotic image group, despite having the same sample size? These scenarios aren’t impossible, but with them coming at a time when Bem was focused on precognition via negative images it’s all quite suspicious.

The Control Isn’t a Control

All too often, researchers using frequentist statistics get blinded by the way p-values ignore the null hypothesis, and don’t bother checking their control groups. Bem’s fairly good about this, but we can do better.

All of Bem’s experiments, save 3 and 4, rely on Bernoulli processes; every person has some probability of guessing the next binary choice correctly, due possibly to inherent precognitive ability, and that probability does not change with time. It follows that the distribution of successful guesses follows the binomial distribution, which can be written:

P( s `divides` p,f ) ~=~ { (s+f)"!" } over { s"!" f"!" } p^s ( 1-p )^f where s is the number of successes, f the number of failures, and p the odds of success; that means P ( s | p,f ) translates to “the probability of having s successes, given the odds of success are p and there were f failures.” Naturally, p must be between 0 and 1.

Let’s try a thought experiment: say you want to test if a single six-sided die is biased to come up 1. You roll it thirty-six times, and observe four instances where it comes up 1. Your friend tosses it seventy-two times, and spots fifteen instances of 1. You’d really like to pool your results together and get a better idea of how fair the die is; how would you do this? If you answered “just add all the successes together, as well as the failures,” you nailed it!The probability distribution of rolling a 1 for a given die, according to you and your friend's experiments.The results look pretty good; both you and your friend would have suspected the die was biased based on your individual rolls, but the combined distribution looks like what you’d expect from a fair die.

But my Bayes 208 post was on conjugate distributions, which defang a lot of the mathematical complexity that comes from Bayesian methods by allowing you to merge statistical distributions. Sit back and think about what just happened: both you and your friend examined the same Bernoulli process, resulting in two experiments and two different binomial distributions. When we combined both experiments, we got back another binomial distribution. The only way this differs from Bayesian conjugate distributions is the labeling; had I declared your binomial to be the prior, and your friend’s to be the likelihood, it’d be obvious the combination was the posterior distribution for the odds of rolling a 1.

Well, almost the only difference. Most sources don’t list the binomial distribution as the conjugate for this situation, but instead the Beta distribution:

Beta( p `divides` %alpha,%beta ) ~=~ { %GAMMA(%alpha + %beta) } over { %GAMMA(%alpha) %GAMMA(%beta) } p^{%alpha-1} ( 1-p )^{%beta-1}

But I think you can work out the two are almost identical, without any help from me. The only real advantage of the Beta distribution is that it allows non-integer successes and failures, thanks to the Gamma function, which in turn permits a nice selection of priors.

In theory, then, it’s dirt easy to do a Bayesian analysis of Bem’s handiwork: tally up the successes and failures from each individual experiment, add them together, and plunk them into a binomial distribution. In practice, there are three hurdles. The easy one is the choice of prior; fortunately, Bem’s datasets are large enough that they swamp any reasonable prior, so I’ll just use the Bayes-Laplace one and be done with it. A bigger one is that we’ve got at least three distinct Bernoulli processes in play: pressing a button to classify an image (experiments 3, 4), remembering a word from a list (8, 9), and guessing the next image out of a binary pair (everything else). If you’re trying to describe precognition and think it varies depending on the input image, then the negative image trials have to be separated from the erotic image ones. Still, this amounts to little more than being careful with the datasets and thinking hard about how a universal precognition would be expressed via those separate processes.

The toughest of the bunch: Bem didn’t record the number of successes and failures, save experiments 8 and 9. Instead, he either saved log timings (experiments 3 and 4) or the success rate, as a percentage of all trials. This is common within frequentist statistics, which is obsessed with maximal likelihoods, but it destroys information we could use to build a posterior distribution. Still, this omission isn’t fatal. We know the number of successes and failures are integer values. If we correctly guess their sum and multiply it by the rate, the result will be an integer; if we pick an incorrect sum, it’ll be a fraction. A complication arrives if there are common factors between the number of successes and the total trials, but there should some results which lack those factors. By comparing results to one another, we should be able to work out both what the underlying total was, as well as when that total changes, and in the process we learn the number of successes and can work backwards to the number of failures.

As the heading suggests, there’s something interesting hidden in the control groups. I’ll start with the binary image pair controls, which behave a lot like a coin flip; as the samples pile up, we’d expect the control distribution to migrate to the 50% line. When we do all the gathering, we find…

What happens when we combine the control groups for the binary image process from Bem (2011).… that’s not good. Experiment 1 had a great control group, but the controls from experiment 5 and 6 are oddly skewed. Since they had a lot more samples, they wind up dominating the posterior distribution and we find ourselves with fully 92.5% of the distribution below the expected value of p = 0.5. This sets up a bad precedent, because we now know that Bem’s methodology can create a skew of 0.67% away from 50%; for comparison, the combined signal from all studies was a skew of 0.83%. Are there bigger skews in the methodology of experiments 2, 3, 4, or 7? We’ve got no idea, because Bem never ran control groups.

Experiments 3 and 4 lack any sort of control, so we’re left to consider the strongest pair of experiments in Bem’s paper, 8 and 9. Bem used a Differential Recall score instead of the raw guess count, as it makes the null effect have an expected value of zero. This Bayesian analysis can cope with a non-zero null, so I’ll just use a conventional success/failure count.

Experiments 8 and 9 from Bem's 2011 paper.

On the surface, everything’s on the up-and-up. The controls have more datapoints between them than the treatment group, but there’s good and consistent separation between them and the treatment. Look very careful at the numbers on the bottom, though; the effects are in quite different places. That’s strange, given the second study only differs from the first via some extra practice (page 14); I can see that improving up the main control and treatment groups, but why does it also drag along the no-practice groups? Either there aren’t enough samples here to get rid of random noise, which seems unlikely, or the methodology changed enough to spoil the replication.

Come to think of it, one of those controls isn’t exactly a control. I’ll let Bem explain the difference.

Participants were first shown a set of words and given a free recall test of those words. They were then given a set of practice exercises on a randomly selected subset of those words. The psi hypothesis was that the practice exercises would retroactively facilitate the recall of those words, and, hence, participants would recall more of the to-be-practiced words than the unpracticed words. […]

Although no control group was needed to test the psi hypothesis in this experiment, we ran 25 control sessions in which the computer again randomly selected a 24-word practice set but did not actually administer the practice exercises. These control sessions were interspersed among the experimental sessions, and the experimenter was uninformed as to condition. [page 13]

So the “no-practice treatment,” as I dubbed it in the charts, is actually a test of precognition! It happens to be a lousy one, as without a round of post-hoc practice to prepare subjects their performance should be poor. Nonetheless, we’d expect it to be as good or better than the matching controls. So why, instead, was it consistently worse? And not just a little worse, either; for experiment 9, it was as worse from its control as the main control was from its treatment group.

What it all Means

I know, I seems to be a touch obsessed with one social science paper. The reason has less to do with the paper than the context around it: you can make a good argument that the current reproducibility crisis is thanks to Bem. Take the words of E.J. Wagenmakers et al.

Instead of revising our beliefs regarding psi, Bem’s research should instead cause us to revise our beliefs on methodology: The field of psychology currently uses methodological and statistical strategies that are too weak, too malleable, and offer far too many opportunities for researchers to befuddle themselves and their peers. […]

We realize that the above flaws are not unique to the experiments reported by Bem (2011). Indeed, many studies in experimental psychology suffer from the same mistakes. However, this state of affairs does not exonerate the Bem experiments. Instead, these experiments highlight the relative ease with which an inventive researcher can produce significant results even when the null hypothesis is true. This evidently poses a significant problem for the field and impedes progress on phenomena that are replicable and important.

Wagenmakers, Eric–Jan, et al. “Why psychologists must change the way they analyze their data: the case of psi: comment on Bem (2011).” (2011): 426.

When it was pointed out Bayesian methods wiped away his results, Bem started doing Bayesian analysis. When others pointed out a meta-analysis could do the same, Bem did that too. You want open data? Bem was a hipster on that front, sharing his data around to interested researchers and now the public. He’s been pushing for replication, too, and in recent years has begun pre-registering studies to stem the garden of forking paths. Bem appears to be following the rules of science, to the letter.

I also know from bitter experience that any sufficiently large research project will run into data quality issues. But, now that I’ve looked at Bem’s raw data, I’m feeling hoodwinked. I expected a few isolated issues, but nothing on this scale. If Bem’s 2011 paper really is a type specimen for what’s wrong with the scientific method, as practiced, then it implies that most scientists are garbage at designing experiments and collecting data.

I’m not sure I can accept that.

How to Become a Radical

If I had a word of the week, it would be “radicalization.” Some of why the term is hot in my circles is due to offline conversations, some of it stems from yet another aggrieved white male engaging in terrorism, and some from yet another study confirms Trump voters were driven by bigotry (via fearing the loss of privilege that comes from giving up your superiority to promote equality).

Some just came in via Rebecca Watson, though, who pointed me to a fascinating study.

For example, a shift from ‘I’ to ‘We’ was found to reflect a change from an individual to a collective identity (…). Social status is also related to the extent to which first person pronouns are used in communication. Low-status individuals use ‘I’ more than high-status individuals (…), while high-status individuals use ‘we’ more often (…). This pattern is observed both in real life and on Internet forums (…). Hence, a shift from “I” to “we” may signal an individual’s identification with the group and a rise in status when becoming an accepted member of the group.

… I think you can guess what Step Two is. Walk away from the screen, find a pen and paper, write down your guess, then read the next paragraph.

The forum investigated here is one of the largest Internet forums in Sweden, called Flashback (…). The forum claims to work for freedom of speech. It has over one million users who, in total, write 15 000 to 20 000 posts every day. It is often criticized for being extreme, for example in being too lenient regarding drug related posts but also for being hostile in allowing denigrating posts toward groups such as immigrants, Jews, Romas, and feminists. The forum has many sub-forums and we investigate one of these, which focuses on immigration issues.

The total text data from the sub-forum consists of 964 Megabytes. The total amount of data includes 700,000 posts from 11th of July, 2004 until 25th of April, 2015.

How did you do? I don’t think you’ll need pen or paper to guess what these scientists saw in Step Three.

We expected and found changes in cues related to group identity formation and intergroup differentiation. Specifically, there was a significant decrease in the use of ‘I’ and a simultaneous increase in the use of ‘we’ and ‘they’. This has previously been related to group identity formation and differentiation to one or more outgroups (…). Increased usage of plural, and decreased frequency of singular, nouns have also been found in both normal, and extremist, group formations (…). There was a decrease in singular pronouns and a relative increase in collective pronouns. The increase in collective pronouns referred both to the ingroup (we) and to one or more outgroups (they). These results suggest a shift toward a collective identity among participants, and a stronger differentiation between the own group and the outgroup(s).

Brilliant! We’ve confirmed one way people become radicalized: by hanging around in forums devoted to “free speech,” the hate dumped on certain groups gradually creates an in-group/out-group dichotomy, bringing out the worst in us.

Unfortunately, there’s a problem with the staircase.

Categories Dictionaries Example words Mean r
Group differentiation First person singular I, my, me -.0103 ***
First person plural We, our, us .0115 ***
Third person plural They, them, their .0081 ***
Certainty Absolutely, sure .0016 NS

***p < .001. NS = not significant. n=11,751.

Table 2 tripped me up, hard. I dropped by the ever-awesome R<-Psychologist and cooked up two versions of the same dataset. One has no correlation, while the other has a correlation coefficient of 0.01. Can you tell me which is which, without resorting to a straight-edge or photo editor?

Comparing two datasets, one with r=0, the other with r=0.01.

I can’t either, because the effect size is waaaaaay too small to be perceptible. That’s a problem, because it can be trivially easy to manufacture a bias at least that large. If we were talking about a system with very tight constraints on its behaviour, like the Higgs Boson, then uncovering 500 bits of evidence over 2,500,000,000,000,000,000 trials could be too much for any bias to manufacture. But this study involves linguistics, which is far less precise than the Standard Model, so I need a solid demonstration of why this study is immune to biases on the scale of r = 0.01.

The authors do try to correct for how p-values exaggerate the evidence in large samples, but they do it by plucking p < 0.001 out of a hat. Not good enough; how does that p-value relate to studies of similar subject matter and methodology? Also, p-values stink. Also also, I notice there’s no control sample here. Do pro-social justice groups exhibit the same trend over time? What about the comment section of sports articles? It’s great that their hypotheses were supported by the data, don’t get me wrong, but it would be better if they’d tried harder to swat down their own hypothesis. I’d also like to point out that none of my complaints falsify their hypotheses, they merely demonstrate that the study falls well short of confirmed or significant, contrary to what I typed earlier.

Alas, I’ve discovered another path towards radicalization: perform honest research about the epistemology behind science. It’ll ruin your ability to read scientific papers, and leave you in despair about the current state of science.

Bayes Bunny iz trying to cool off after reading too many scientific papers.

The Laziness of Steven Pinker

I know, I know, I should have promoted that OrbitCon talk on Steven Pinker before it aired. I was a bit swamped developing material for it, ironically, most of which never made it to air. Don’t worry, I’ll be sharing the good bits via blog post. Amusingly, this first example isn’t from that material. I wound up reading a lot of Pinker, and developed a hunch I wasn’t able to track down before air time. In a stroke of luck, Siggy handed me the material I needed to properly follow up.

Enough suspense: what’s your opinion of self-plagiarism, or copying your own work without flagging what you’ve done?

… self-plagiarism does carry with it some level of dishonesty, at least in some situations. The problem is that, when an author, artist or other creator presents a new work, it’s generally expected to be all-new content, unless otherwise clearly stated. … with an academic paper, one is generally expected to showcase what they have learned most recently, meaning that self-plagiarism defeats the purpose of the paper or the assignment. On the other hand, in a creative environment, however, reusing old passages, especially in a limited manner, might be more about homage and maintaining consistency than plagiarism.

It’s a bit of a gray area, isn’t it? The US Office of Research Integrity declares it unethical, but also declares that self-plagiarism isn’t misconduct. Nonetheless it could be considered misconduct in an academic context, and the ORI themselves outline the case:

For example, in one editorial, Schein (2001) describes the results of a study he and a colleague carried out which found that 92 out of 660 studies taken from 3 major surgical journals were actual cases of redundant publication. The rate of duplication in the rest of the biomedical literature has been estimated to be between 10% to 20% (Jefferson, 1998), though one review of the literature suggests the more conservative figure of approximately 10% (Steneck, 2000). However, the true rate may depend on the discipline and even the journal and more recent studies in individual biomedical journals do show rates ranging from as low as just over 1% in one journal to as high as 28% in another (see Kim, Bae, Hahm, & Cho, 2014) The current situation has become serious enough that biomedical journal editors consider redundancy and duplication one of the top areas of concern (Wager, Fiack, Graf, Robinson, & Rowlands, 2009) and it is the second highest cause for articles to be retracted from the literature between the years 2007 and 2011 (Fang, Steen, & Casadevall, 2012).

But is it misconduct in the context of non-academic science writing? I’m not sure, but I think it’s fair to say self-plagiarism counts as lazy writing. Whatever the ethics, let’s examine an essay by Pinker that Edge published sometime before January 10th, 2017, and match it up against Chapter 2 of Enlightenment Now. I’ve checked the footnotes and preface of the latter, and failed to find any reference to that Edge essay, while the former does not say it’s excerpted from a forthcoming book. You’d have no idea one copy existed if you’d only read the other, so any matching passages count as self-plagiarism.

How many passages match? I’ll use the Edge essay as a base, and highlight exact duplicates in red, sections only present in Enlightenment Now in green, paraphrases in yellow, and essay-only text in black.

The Second Law of Thermodynamics states that in an isolated system (one that is not taking in energy), entropy never decreases. (The First Law is that energy is conserved; the Third, that a temperature of absolute zero is unreachable.) Closed systems inexorably become less structured, less organized, less able to accomplish interesting and useful outcomes, until they slide into an equilibrium of gray, tepid, homogeneous monotony and stay there.

In its original formulation the Second Law referred to the process in which usable energy in the form of a difference in temperature between two bodies is inevitably dissipated as heat flows from the warmer to the cooler body. (As the musical team Flanders & Swann explained, “You can’t pass heat from the cooler to the hotter; Try it if you like but you far better notter.”) A cup of coffee, unless it is placed on a plugged-in hot plate, will cool down. When the coal feeding a steam engine is used up, the cooled-off steam on one side of the piston can no longer budge it because the warmed-up steam and air on the other side are pushing back just as hard.

Once it was appreciated that heat is not an invisible fluid but the energy in moving molecules, and that a difference in temperature between two bodies consists of a difference in the average speeds of those molecules, a more general, statistical version of the concept of entropy and the Second Law took shape. Now order could be characterized in terms of the set of all microscopically distinct states of a system (in the original example involving heat, the possible speeds and positions of all the molecules in the two bodies). Of all these states, the ones that we find useful from a bird’s-eye view (such as one body being hotter than the other, which translates into the average speed of the molecules in one body being higher than the average speed in the other) make up a tiny sliver of the possibilities, while the disorderly or useless states (the ones without a temperature difference, in which the average speeds in the two bodies are the same) make up the vast majority. It follows that any perturbation of the system, whether it is a random jiggling of its parts or a whack from the outside, will, by the laws of probability, nudge the system toward disorder or uselessness —not because nature strives for disorder, but because there are so many more ways of being disorderly than of being orderly. If you walk away from a sand castle, it won’t be there tomorrow, because as the wind, waves, seagulls, and small children push the grains of sand around, they’re more likely to arrange them into one of the vast number of configurations that don’t look like a castle than into the tiny few that do. [Enlightenment Now adds five sentences here.]

 

I could (and have!) carried on, demonstrating that almost all of that essay reappears in Pinker’s book. Maybe half of the reappearance is verbatim. I figure he copy-pasted the contents of his January 2017 essay into the manuscript for his 2018 book, and expanded it to fill an entire chapter. Whether I’m right or wrong, I think the similarities make a damning case for intellectual laziness. It also sets up a bad precedent: if Pinker can get this lazy with his non-academic writing, how lazy can he be with his academic work? I haven’t looked into that, and I’m curious if anyone else has.

The Tuskegee Syphilis Study

Was it three years ago? Almost to the day, from the looks of it.

Biomedical research, then, promises vast increases in life, health, and flourishing. Just imagine how much happier you would be if a prematurely deceased loved one were alive, or a debilitated one were vigorous — and multiply that good by several billion, in perpetuity. Given this potential bonanza, the primary moral goal for today’s bioethics can be summarized in a single sentence.

Get out of the way.

A truly ethical bioethics should not bog down research in red tape, moratoria, or threats of prosecution based on nebulous but sweeping principles such as “dignity,” “sacredness,” or “social justice.” Nor should it thwart research that has likely benefits now or in the near future by sowing panic about speculative harms in the distant future.

That was Steven Pinker arguing that biomedical research is too ethical. Follow that link and you’ll see my counter-example: the Tuskegee syphilis study. It is a literal textbook example of what not to do in science. Pinker didn’t mention it back then, but it was inevitable he’d have to deal with it at some time. Thanks to PZ, I now know he has.

At a recent conference, another colleague summed up what she thought was a mixed legacy of science: vaccines for smallpox on the one hand; the Tuskegee syphilis study on the other. In that affair, another bloody shirt ind the standard narrative about the evils of science, public health researchers, beginning in 1932, tracked the progression of untreated latent syphilis in a sample of impoverished African Americans for four decades. The study was patently unethical by today’s standards, though it’s often misreported to pile up the indictment. The researchers, many of them African American or advocates of African American health and well-being, did not infect the participants as many people believe (a misconception that has led to the widespread conspiracy theory that AIDS was invented in US government labs to control the black population). And when the study began, it may even have been defensible by the standards of the day: treatments for syphilis (mainly arsenic) were toxic and ineffective; when antibiotics became available later, their safety and efficacy in treating syphilis were unknown; and latent syphilis was known to often resolve itself without treatment. But the point is that the entire equation is morally obtuse, showing the power of Second Culture talking points to scramble a sense of proportionality. My colleague’s comparison assumed that the Tuskegee study was an unavoidable part of scientific practice as opposed to a universally deplored breach, and it equated a one-time failure to prevent harm to a few dozen people with the prevention of hundreds of millions of deaths per century in perpetuity.

What horse shit.

To persuade the community to support the experiment, one of the original doctors admitted it “was necessary to carry on this study under the guise of a demonstration and provide treatment.” At first, the men were prescribed the syphilis remedies of the day — bismuth, neoarsphenamine, and mercury — but in such small amounts that only 3 percent showed any improvement. These token doses of medicine were good public relations and did not interfere with the true aims of the study. Eventually, all syphilis treatment was replaced with “pink medicine” — aspirin. To ensure that the men would show up for a painful and potentially dangerous spinal tap, the PHS doctors misled them with a letter full of promotional hype: “Last Chance for Special Free Treatment.” The fact that autopsies would eventually be required was also concealed. As a doctor explained, “If the colored population becomes aware that accepting free hospital care means a post-mortem, every darky will leave Macon County…”

  • “it equated a one-time failure to prevent harm to a few dozen people”: In reality, according to that last source, “28 of the men had died directly of syphilis, 100 were dead of related complications, 40 of their wives had been infected, and 19 of their children had been born with congenital syphilis.” As of August last year, 12 former children were still receiving financial compensation.
  • “the prevention of hundreds of millions of deaths per century in perpetuity”: In reality, the Tuskegee study wasn’t the only scientific study looking at syphilis. Nor even the first. Syphilis was discovered in 1494, named in 1530, the causative organism was found in 1905, and the first treatments were developed in 1910. The science was dubious at best:

The study was invalid from the very beginning, for many of the men had at one time or another received some (though probably inadequate) courses of arsenic, bismuth and mercury, the drugs of choice until the discovery of penicillin, and they could not be considered untreated. Much later, when penicillin and other powerful antibiotics became available, the study directors tried to prevent any physician in the area from treating the subjects – in direct opposition to the Henderson Act of 1943, which required treatment of venereal diseases.

A classic study of untreated syphilis had been completed years earlier in Oslo. Why try to repeat it? Because the physicians who initiated the Tuskegee study were determined to prove that syphilis was ”different” in blacks. In a series of internal reviews, the last done as recently as 1969, the directors spoke of a ”moral obligation” to continue the study. From the very beginning, no mention was made of a moral obligation to treat the sick.

Pinker’s response to the Tuskegee study is to re-write history to suit his narrative, again. No wonder he isn’t a fan of ethics.

Steven Pinker, “Historian”

It’s funny, if you look back over my blog posts on Steven Pinker, you’ll notice a progression.

Ignoring social justice concerns in biomedical research led to things like the Tuskegee experiment. The scientific establishment has since tried to correct that by making it a critical part. Pinker would be wise to study the history a bit more carefully, here.


Setting aside your ignorance of the evidence for undercounting in the FBI’s data, you can look at your own graph and see a decline?

When Sargon of Arkkad tried and failed to discuss sexual assault statistics, he at least had the excuse of never having gotten a higher education, never studying up on the social sciences. I wonder what Steven Pinker’s excuse is.


Ooooh, I get it. This essay is just an excuse for Pinker to whine about progressives who want to improve other people’s lives. He thought he could hide his complaints behind science, to make them look more digestible to himself and others, but in reality just demonstrated he understands physics worse than most creationists. What a crank.

You’ll also notice a bit of a pattern, too, one that apparently carries on into Pinker’s book about the Enlightenment.

It is curious, then, to find Pinker breezily insisting that Enlightenment thinkers used reason to repudiate a belief in an anthropomorphic God and sought a “secular foundation for morality.” Locke clearly represents the opposite impulse (leaving aside the question of whether anyone in period believed in a strictly anthropomorphic deity).

So, too, Kant. While the Prussian philosopher certainly had little use for the traditional arguments for God’s existence – neither did the exceptionally pious Blaise Pascal, if it comes to that – this was because Kant regarded them as stretching reason beyond its proper limits. Nevertheless, practical reason requires belief in God, immorality and a post-mortem existence that offers some recompense for injustices suffered in the present world.

That’s from Peter Harrison, a professional historian. Even I was aware of this, though I am guilty of a lie of omission. I’ve brought up the “Cult of Reason” before, which was a pseudo-cult set up during the French Revolution that sought to tear down religion and instead worship logic and reason. What I didn’t mention was that it didn’t last long; Robespierre shortly announced his “Cult of the Supreme Being,” which promoted Deism as the official religion of France, and had the leaders of the Cult of Reason put to death. Robespierre himself was executed shortly thereafter, for sounding too much like a dictator, and after a half-hearted attempt at democracy France finally settled on Napoleon Bonaparte, a dictator everyone could get behind. The shift to reason and objectivity I was hinting at back then was more gradual than I implied.

If we go back to the beginning of the scientific revolution – which Pinker routinely conflates with the Enlightenment – we find the seminal figure Francis Bacon observing that “the human intellect left to its own course is not to be trusted.” Following in his wake, leading experimentalists of the seventeenth century explicitly distinguished what they were doing from rational speculation, which they regarded as the primary source of error in the natural sciences.

In the next century, David Hume, prominent in the Scottish Enlightenment, famously observed that “reason alone can never produce any action … Reason is, and ought only to be the slave of the passions.” And the most celebrated work of Immanuel Kant, whom Pinker rightly regards as emblematic of the Enlightenment, is the Critique of Pure Reason. The clue is in the title.

Reason does figure centrally in discussions of the period, but primarily as an object of critique. Establishing what it was, and its intrinsic limits, was the main game. […]

To return to the general point, contra Pinker, many Enlightenment figures were not interested in undermining traditional religious ideas – God, the immortal soul, morality, the compatibility of faith and reason – but rather in providing them with a more secure foundation. Few would recognise his tendentious alignment of science with reason, his prioritization of scientific over all other forms of knowledge, and his positing of an opposition between science and religion.

I’m just skimming Harrison’s treatment, the rest of the article is worth a detour, but it really helps underscore how badly Pinker wants to re-write history. Here’s something the man himself committed to electrons:

More insidious than the ferreting out of ever more cryptic forms of racism and sexism is a demonization campaign that impugns science (together with the rest of the Enlightenment) for crimes that are as old as civilization, including racism, slavery, conquest, and genocide. […]

“Scientific racism,” the theory that races fall into a hierarchy of mental sophistication with Northern Europeans at the top, is a prime example. It was popular in the decades flanking the turn of the 20th century, apparently supported by craniometry and mental testing, before being discredited in the middle of the 20th century by better science and by the horrors of Nazism. Yet to pin ideological racism on science, in particular on the theory of evolution, is bad intellectual history. Racist beliefs have been omnipresent across history and regions of the world. Slavery has been practiced by every major civilization and was commonly rationalized by the belief that enslaved peoples were inherently suited to servitude, often by God’s design. Statements from ancient Greek and medieval Arab writers about the biological inferiority of Africans would curdle your blood, and Cicero’s opinion of Britons was not much more charitable.

More to the point, the intellectualized racism that infected the West in the 19th century was the brainchild not of science but of the humanities: history, philology, classics, and mythology.

As I’ve touched on, this is so far from reality it’s practically creationist. Let’s ignore the implication that no-one used science to promote racism past the 1950’s, which ain’t so, and dig up more data points on the dark side of the Enlightenment.

… the Scottish philosopher David Hume would write: “I am apt to suspect the Negroes, and in general all other species of men to be naturally inferior to the whites. There never was any civilized nation of any other complection than white, nor even any individual eminent in action or speculation.” […]

Another two decades on, Immanuel Kant, considered by many to be the greatest philosopher of the modern period, would manage to let slip what is surely the greatest non-sequitur in the history of philosophy: describing a report of something seemingly intelligent that had once been said by an African, Kant dismisses it on the grounds that “this fellow was quite black from head to toe, a clear proof that what he said was stupid.” […]

Scholars have been aware for a long time of the curious paradox of Enlightenment thought, that the supposedly universal aspiration to liberty, equality and fraternity in fact only operated within a very circumscribed universe. Equality was only ever conceived as equality among people presumed in advance to be equal, and if some person or group fell by definition outside of the circle of equality, then it was no failure to live up to this political ideal to treat them as unequal.

It would take explicitly counter-Enlightenment thinkers in the 18th century, such as Johann Gottfried Herder, to formulate anti-racist views of human diversity. In response to Kant and other contemporaries who were positively obsessed with finding a scientific explanation for the causes of black skin, Herder pointed out that there is nothing inherently more in need of explanation here than in the case of white skin: it is an analytic mistake to presume that whiteness amounts to the default setting, so to speak, of the human species.


Indeed, connections between science and the slave trade ran deep during [Robert] Boyle’s time—all the way into the account books. Royal Society accounts for the 1680s and 1690s shows semi-regular dividends paid out on the Society’s holdings in Royal African Company stock. £21 here, £21 there, once a year, once every two years. Along with membership dues and the occasional book sale, these dividends supported the Royal Society during its early years.

Boyle’s early “experiments” with the inheritance of skin color set an agenda that the scientists of the Royal Society pursued through the decades. They debated the origins of blackness with rough disregard for the humanity of enslaved persons even as they used the Royal African’s Company’s dividends to build up the Royal Society as an institution. When it came to understanding skin color, Boyle used his wealth and position to help construct a science of race that, for centuries, was used to justify the enslavement of Africans and their descendants globally.


This timeline gives an overview of scientific racism throughout the world, placing the Eugenics Record Office within a broader historical framework extending from Enlightenment-Era Europe to present-day social thought.

All this is obvious via a glance at a history book, something which Pinker is apparently allergic to. I’ll give Harrison the final word:

If we put into the practice the counting and gathering of data that Pinker so enthusiastically recommends and apply them to his own book, the picture is revealing. Locke receives a meagre two mentions in passing. Voltaire clocks up a modest six references with Spinoza coming in at a dozen. Kant does best of all, with a grand total of twenty-five (including references). Astonishingly, Diderot rates only two mentions (again in passing) and D’Alembert does not trouble the scorers. Most of these mentions occur in long lists. Pinker refers to himself over 180 times. […]

… if Enlightenment Now is a model of what Pinker’s advice to humanities scholars looks like when put into practice, I’m happy to keep ignoring it.

Computational Propaganda

Sick of all this memo talk? Too bad, because thanks to Lynna, OM in the Political Madness thread I discovered a new term: “computational propaganda,” or the use of computers to help spread talking points and generate “grassroots” activism. It’s a lot more advanced than running a few bots, too. You’ll have to read the article to learn the how and why, but I can entice you with its conclusion:

The problem with the term “fake news” is that it is completely wrong, denoting a passive intention. What is happening on social media is very real; it is not passive; and it is information warfare. There is very little argument among analytical academics about the overall impact of “political bots” that seek to influence how we think, evaluate and make decisions about the direction of our countries and who can best lead us—even if there is still difficulty in distinguishing whose disinformation is whose. Samantha Bradshaw, a researcher with Oxford University’s Computational Propaganda Research Project who has helped to document the impact of “polbot” activity, told me: “Often, it’s hard to tell where a particular story comes from. Alt-right groups and Russian disinformation campaigns are often indistinguishable since their goals often overlap. But what really matters is the tools that these groups use to achieve their goals: Computational propaganda serves to distort the political process and amplify fringe views in ways that no previous communication technology could.”

This machinery of information warfare remains within social media’s architecture. The challenge we still have in unraveling what happened in 2016 is how hard it is to pry the Russian components apart from those built by the far- and alt-right—they flex and fight together, and that alone should tell us something. As should the fact that there is a lesser far-left architecture that is coming into its own as part of this machine. And they all play into the same destructive narrative against the American mind.

Democracies have not faced a challenge like this since yellow journalism.

Winning Hearts and Minds

I’ll forgive you if haven’t heard of Ian Danskin, if only because he’s primarily known on YouTube as Innuendo Studios. You know, the person behind “Why Are You So Angry?” and more recently “The Alt-Right Playbook.” The latter project is aimed at sharpening the rhetoric of progressives to better defend against the “playbook” the Alt-Right uses in online arguments. It’s still a work in progress, but recently Danskin tried to jump ahead and compress it all into a single lecture.

[Read more…]

The Power Of Representation

I kicked around a number of titles for this one: The Persistence of Bias, Science is Social, Beeing Blind. It’s amazing just how many themes can be packed into a Twitter thread.

Hank Campbell: Resist the call to make science about social justice. Astronomers should not be enthusiastic when told that their cosmic observations are inevitably a reflex of the power of the socially privileged.

Ask An Entomologist: Although we disagree with this tweet…it gives us an opportunity to explore a really interesting topic. What we now call ‘queen’ bees-the main female reproductive honeybees-were erroneously called ‘kings’ for nearly 2,000 years. Why? Let’s explore the history of bees!

We’ve been keeping bees for 5,000 years+ and what we called the various classes of bees was closely tied to the societies naming those classes. For instance, in a lot of societies it was very common to call the ‘workers’ slaves because slavery was common at the time. For awhile, this was the big head-honcho in the biological sciences. This is Aristotle, whose book The History of Animals was the accepted word on animal biology in Europe until roughly the 1600s. This book was published in 350, and discussed honeybees in quite some detail …and is a good reflection on what was known at the time. […] I’d recommend reading the whole thing…it’s really interesting for a number of reasons.

…but in particular, let’s look at how Aristotle described the swarming process. Bees reproduce by swarming: They make new queens, who leave to set up a new hive. The queens take a big chunk of the colony’s workers with them.

“Of the king bees there are, as has been stated, two kinds. In every hive there are more kings than one; and a hive goes to ruin if there be too few kings, not because of anarchy thereby ensuing, but, as we are told, because these creatures contribute in some way to the generation of the common bees. A hive will go also to ruin if there be too large a number of kings in it; for the members of the hives are thereby subdivided into too many separate factions.”

Aristotle didn’t know what we know about bees now…but it was widely accepted that the biggest bees in the colony lead the hive somehow and were essential for reproduction and swarming. …but we now know the queens are female. Why didn’t Aristotle?

Well it turns out that Aristotle, frankly, had some *opinions* about women. He was…uh, a little sexist. Which was, like, common at the time. Without going into all of his views on the topic, it’s apparent his views on women pretty heavily influenced what he saw was going on in the beehive. He thought of reproduction as a masculine activity, and thought of women as property. He…just wasn’t very objective about this. So, when he saw a society led almost entirely of women…it actually makes a lot of sense as to why he saw the ‘queen’ bees as male and called them kings. These ideas of women in his circle were so ingrained that a female ruler literally wouldn’t compute.

Moving on through the middle ages, the name ‘king’ kind of stuck because biological sciences were stuck on Aristotle’s ideas for a very long time. Beekeepers *knew* the queens were female; they were observed laying eggs…but their exact role was controversial outside of them. In fact, in most circles, it was commonly accepted that the workers gathered the larvae which grew on plants. Again, this is from Aristotle’s work.

So…today it’s completely and 100% accepted that queen bees are, in fact, female…and that the honeybee society is led by women. What changed in Western Society to get this idea accepted?

The exact work which popularized the (scientifically accurate) idea of the honeybee as a female-led society was The Feminine Monarchy, by Charles Butler. However, I’d argue this lady also played a role. The woman in the picture … is Queen Elizabeth, who ruled England from 1558 until her death in 1603. Charles Butler (1560-1647) published The Feminine Monarchy in 1609, and had lived under Queen Elizabeth’s reign for most of his life. This is largely a ‘right place, right time’ situation. At this point, there was a lot of science that was just up and starting. There had been female rulers before, but not at the exact point where people were rethinking their assumptions. The fact that Charles Butler was interested in bees, *and* lived under a female monarch for most of his life, I think played a major role in his decision to substitute one simple word in his book.

That substitution? He called ‘king bees’ ‘queen bees’…and it stuck.

At this point in Europe’s history, there had been several female monarchs so the idea of a female leader didn’t seem so odd. Society was simply primed to accept the idea of a female ruler.

…but this thread isn’t just about words, it’s also about *sex*.

How so? Sorry, you’ll have to click through for that one. Bee sure to read to the end for the punchline, too. Big kudos are due to @BugQuestions for such an expansive, deep Twitter thread.

The Gender Inclusivity of Diverse White Privilege Equity

Blame Shiv for this one; she posted about someone at Monsanto inviting Jordan Peterson to talk about GMOs, and it led me down an interesting rabbit hole. For one thing, the event already happened, and it was the farce you were expecting. This, however, caught my eye:

Corrupt universities—and Women’s Studies departments in particular, he says—are responsible for turning students into activists who will one day tear apart the fabric of society. “The world runs on ideas. And the ideas that are in the universities are the ideas that are going to be in the general public in five to ten years. And there’s no shielding yourself from it,” he said.

Peterson also shared a trick for figuring out whether or not a child’s school has been affected by the coming crisis: If a schoolteacher uses any of the five words listed on his display screen—”diversity,” “inclusivity,” “equity,” “white privilege,” or “gender”—then a child has been “exposed.”

What’s Peterson’s solution for all this? “The answer to the ills that our society still obviously suffers from,” he said, is that “people should adopt an ethos of responsibility rather than continually clamoring about their rights, which is something that we’ve been talking about for about four decades too long, as far as I can surmise.”

Four decades puts us back into the 1970’s, when women’s liberation groups were calling to be able to exercise their right to bodily autonomy, to be free from violence, to equal pay for equal work, to equal custody of kids. If Peterson is opposed to that then he’s more radical than most MRAs, who are generally fine with Second Wave feminism. I wonder if he’s a lost son of Phyllis Schlafly.

But more importantly, he appears to be warning us of a crisis coming in 5-10 years, one that invokes those five terms as holy writ. That’s …. well, let’s step through it.

[Read more…]