Sunday, December 16, 2012

Implicature and the Interpretation of the Law (Part One)

Paul Grice, originator of the theory of conversational implicature


Consider the following example (lifted unashamedly from Steven Pinker’s book The Stuff of Thought):

A gangster walks into a local restaurant. The restaurant has been doing well recently, and the local criminal gangs are aware of this fact. The gangster walks over to the restaurant owner, stares conspicuously around the room, and says “This is real nice place you got here. It would be a shame if something happened to it.”

Ostensibly, the gangster’s statement is one of fact: depending on what the “something” in question is, it may indeed be a shame if it happened to the restaurant. But of course no one reading the statement really thinks it is as innocuous as that. Everyone knows that it constitutes a thinly-veiled threat. Why is this?

The answer lies in something known as conversational implicature, which is the fancy label given to the mundane phenomenon that the semantic content of a particular utterance or sentence is not exhausted by the meaning of the words that make up that utterance. Which is to say: it is possible for an utterance to have an implied meaning, which is just as important and just as readily understood as that of its explicit meaning. Indeed, sometimes it is more important than the explicit meaning, as in the case of the gangster’s veiled threat: If the restaurant owner didn’t pick up on the implied meaning, he could create problems for himself.

While the phenomenon of implicature is mundane, it can lead to problems in particular contexts. One of those contexts is the law. In a certain sense, laws are created through speech acts. Legislatures and legal officials “speak” the law in the form of both written and oral utterances. Is it possible for those utterances to imply more than they explicitly say? And if so, is it acceptable for judges to appeal to those implied meanings when interpreting and applying the law?

Over the next two posts I want to look at these questions, and I do so with the help of Francesca Poggi’s article “Law and Conversational Implicatures” (which appears in the impressively obscure International Journal of Semiotics and Law). In this post, I kick things off by outlining Grice’s classic theory of conversational implicature, before then considering the distinction between generalised and particularised implicatures. In the next post, I’ll address the application of these concepts to the law. As we’ll see, Poggi thinks that implicature has a limited role to play in the interpretation of statutes and other “authoritative legal acts”, but it could have a more expansive role to play in the interpretation of contracts and other “private acts of autonomy”.


1. Grice on Conversational Implicature
The classic model for understanding how conversational implicature works was developed by the philosopher Paul Grice (pictured above). His model is built around something he calls the cooperative principle. This principle allegedly governs most ordinary conversational exchanges, and is constituted by a number of maxims. Let’s work our way through Grice’s account in a bit more detail.

Let’s start with a model utterance which will illustrate the phenomenon of implicature:

(a) “I am reading John’s book.

Although perfectly natural as a linguistic construct, this utterance is ambiguous. If we focused purely on the semantic content of the words that it contains, we would be left with at least two plausible interpretations. Either I am saying that I am reading a book that was written by John, or I am saying that I am reading a book that is owned or possessed by John. Nothing in the words tells us which of the two meanings should apply.

If the utterance really is ambiguous in this manner, then one is left with the burning question: why say things this way? Why is an utterance like this perfectly natural even though it has two possible meanings? The answer lies in the cooperative principle. According to Grice, in ordinary conversational exchanges, we all tend to adhere to the following principle:

Cooperative Principle (CP): Make your conversational contribution such as is required, at the stage at which it occurs, by the accepted purpose or direction of the talk exchange in which you are engaged.

In essence, the cooperative principle holds that whenever you make a contribution to a conversation, you should say whatever is required to convey your intended meaning, but no more than is required. In other words, if the context in which the conversation takes place makes it clear that utterance (a) is a reference to a book that John has written, then utterance (a) is the acceptable way in which to convey that meaning, despite the latent linguistic ambiguity. The participants in the conversation will be able to work out the implication for themselves; no more needs to be said. For example, suppose we are attending a book launch, to celebrate John’s recently published book. You find me thumbing through the pages of a book, and ask me what I am reading. I reply by saying “I am reading John’s book”. In this context, it’s perfectly clear which of the two possible meanings applies.

Grice unpacks the cooperative principle by breaking it down into a series of maxims. They are as follows:

Maxims of Quantity:
Be as informative as is required. But be no more informative than is required.
Maxims of Quality:
Do not say what you believe to be false. Do not say that for which you lack evidence.
Maxim of Relation:
Be relevant.
Maxims of Manner:
Be clear, avoid obscurity and ambiguity. Be brief and be orderly.

Now some of these maxims seem a little unhelpful, particularly those counseling against ambiguity, since the phenomenon of implicature, at least as illustrated by the example of utterance (a), seems arise even though they are violated. But in many ways that’s the whole point. As Poggi notes, implicature depends both on the meaning of the words used and on the maxims that apply to the particular conversational context. But which maxims apply in which context is variable. Thus, in some contexts the avoidance of ambiguity is trumped by the efficiency and brevity of communication.

A good example of this is sarcasm. If I say to you “that was a real funny joke”, it’s likely that I’m being sarcastic. This will usually be obvious, thanks to both the context (no one laughed) and the manner of speech (inflection and tone). This is a classic example of implicature, since the implied meaning of what I say diverges considerably (indeed, orthogonally) from the linguistic meaning of what I say. But this is made possible by the deliberate and obvious violation of the first maxim of quality: do not say what you believe to be false. We both know that this maxim usually applies to our conversations, but in this context its deliberate violation creates a dramatic effect without leading to any confusion about the intended meaning.

All of which leaves us wondering about the precise the status of the cooperative principle and the associated maxims. Are they prescriptive? In other words, should we follow them? Or are they descriptive? Do they merely describe what is typically happening when people communicate successfully?

Poggi opts for a quasi-prescriptive interpretation of the principle and the maxims. She views them as “customary hermeneutical technical rules”, which means:

Poggi’s Rule: If you follow the CP and its associated maxims then you will (in general) cooperate, understand what others are saying, and be understood.

The “in general” clause is key here (and is my addition) since, as we have just seen, it is possible to be understood even when you do violate the maxims. But that is only because (and if) the context makes clear what the implicature really is. If I send you a text message saying “I am reading John’s book”, and there is no preceding context in which my utterance is situated, then ambiguity becomes a problem. It’s highly likely that you’ll need to ask me to clarify the intended meaning. All of which brings us to the next issue: the distinction between particularised and generalised implicatures.


2. General and Particularised Implicatures
The basic idea of implicature is straightforward: utterances can often mean more than what they say. But its manifestations are many and complex. One of the complexities arises from the fact that there can be generalised and particularised implicatures. That is to say: implicatures that hold true across all contexts, and implicatures that only arise in specific contexts. Here’s an example of the former:

(b) “I went into a house”

This carries the general implicature that the house was not mine. Thus, the utterance could be construed as “I went into a house and the house was not mine”, but the italicised portion is left unsaid. The reason being that referring to the house using the indefinite article is generally understood as being the way to refer to a house that is not yours. The normal way of referring to one’s own house would be to say “I went into my house”.

According to Poggi, generalised implicatures are made possible by the maxim of quantity — one says no more than needs to be said — and are partially (if not entirely) independent of the speaker’s intentions. In other words, the implicature arises even if the speaker did not directly intend it. This could actually be a problem in some instances, for a sentence could carry a generalised implicature that actually defeats the speaker’s intentions. For example, I could say “I went into a house” and intend for it to be understood that the house was mine, but unfortunately listeners would not pick up on this due to the generalised implicature. This, however, would be my fault since I chose an inappropriate string of words to convey my intended meaning.

The situation is very different when it comes to particularised implicatures. These only arise in a specific context, and they only work when that context is shared by both the speaker and the listener. Consider the following conversation in a restaurant after the bill has been paid:

(c) Andy: “I’m sorry I made Paul pay the bill.” 
     Barry: “Paul owns four houses.”

Here, the implied meaning of Barry’s utterance is that Andy should not feel sorry for Paul, since Paul owns four houses and is thus wealthy enough to pay for the meal. The implied meaning is understood by both the speaker and the listener in the specific context. But if we detached Barry’s utterance from the specific context, no such implicature would arise.

Furthermore, Barry’s implied meaning might not be appreciated by Andy if the background context of the conversation is not fully shared. For instance, Andy might know that Paul is in a lot of financial trouble because of his properties, but Barry might not. Thus, Andy might think that Barry is merely emphasising his own remorse by highlighting the financial troubles. But Barry might intend the exact opposite because he knows nothing of the financial troubles. In this instance, there is a communication failure, and it is attributable to the lack of a shared context. This is a significant point, and one we shall return to in part two when discussing implicature in the law.

All of which brings us back to the example at the start of this blog post. As we saw, when the gangster says to the restaurant owner “this is a real nice place, it would be a shame if something happened to it”, the implication is that this is a threat. But is this implication generalised or particularised? The obvious answer is to say that it is particularised. After all, detached from the story of the gangster and the restaurant owner, that string of words carries with it no obvious implication.

Or does it? This is an interesting case. If I saw those words strung together in that particular order, but detached from a specific conversational context, I would still be inclined to think they contained an implied threat. This is because this linguistic form of the implied threat is so common in popular culture. Thus, it might be that the gangster’s utterance has a generalised implicature. Pinker points to a similar phenomenon in relation to the request “Would you like to come up and see my etchings?”, which, in our culture, is almost always understood to imply an invitation to sexual congress. These examples suggest that the line between the particularised and the generalised implicature might be a fuzzy and somewhat fluid one. Something could start out life as a particularised implicature, but if it becomes widely known, it may end up a generalised implicature.

Anyway, we shall leave it there for now. As we have seen, utterances often contain implicatures. That is: they imply more than they actually say. This is made possible, according to Grice, by the cooperative principle of ordinary conversation, and its associated maxims, although the applicability of these maxims may vary depending on the context. Furthermore, implicatures can be generalised or particularised. If generalised, they always arise whenever the relevant utterance is made. If particularised, they only arise in a specific context, provided that the context is shared by both the speaker and the listener. This raises all sorts of interesting questions for the law. We’ll look at these in part two.

Monday, December 10, 2012

Schauer on (fMRI) Lie Detection in the Law (Part Two)



(Part One)

This is the second part in short series of posts looking at Frederick Schauer’s article “Lie Detection, Neuroscience and the Law of Evidence”. In this article, Schauer examines the debate surrounding the legal admissibility of fMRI lie detection evidence, and argues that there are good reasons to allow such evidence in a court of law. This is interesting in that it runs contrary to the prevailing view about fMRI lie detection.

In part one, I reviewed some of the background issues in Schauer’s article. This included a brief discussion of the problem of false testimony within the law — a problem that makes a reliable lie detector particularly alluring. It also included an overview of the legal history of the lie detector test, noting that from its earliest days it has struggled to win acceptance in the courts. This trend has continued despite the advent of newer versions of the test using fMRI imaging techniques.

Schauer questions the tenability of this trend. He does so by defending one overarching claim, which we may call “Schauer’s Thesis”:

Schauer’s Thesis: Whether fMRI lie detection evidence should be admitted to court is not simply a question of its scientific validity and reliability, it also (perhaps primarily) a question of the normative and ethical function of the law. That is to say, questions of evidential admissibility are fundamentally determined by legal-ethical standards, not purely scientific ones.

This claim is significant in that current tests for the admissibility of scientific evidence, such as DNA fingerprinting and other forensic techniques, are heavily reliant on scientific standards of validity and reliability. For instance the Daubert test, which is now advocated for introduction in the UK, states that judges should assess scientific evidence by referring to the various indicia of reliability that are common in the scientific world. These indicia include things like “known error rates”, “general acceptance within the relevant scientific community”, “testability” and “passing peer review”.

This approach yields significant legal territory to the norms of scientific inquiry, and while this may often be appropriate, Schauer’s Thesis urges lawyers and legal theorists to regain at least part of this territory. What scientists rightfully deem “good evidence” and what legal theorists rightfully deem “good evidence” may be two different things. It’s important not to lose sight of this.

Schauer supports his thesis with two arguments. For cognitive convenience, I have labelled them the probative context argument and the epistemic progress argument. In the remainder of this post, I examine each argument in some detail.


1. The Probative Context Argument
As mentioned in part one, in every legal case there is some set of facts that need to be proved (or disproved) in order for the case to succeed. If I am to be convicted of murder, it must be proved that I intentionally killed another person. One step on the way towards proving this would be to establish that I was present at the scene of the crime. Lie detectors or other forensic evidence might be used to do this. But the value of any such evidence depends largely on three factors:

Probability: Does the evidence raise or lower the probability of the factum probandum and if so, by how much does it raise or lower its probability?
Standard of Proof: What confidence threshold must the probability of the factum probandum cross in order for it to count as being proved or not proved?
Legal Purpose: Is the evidence being submitted in order to prove or disprove the factum probandum?

These three factors determine the probative context in which the evidence is presented. This context varies relative to the legal issue at stake, and the party for whom the evidence is proffered.

For example, in criminal cases, the standard of proof for the prosecution is beyond a reasonable doubt. This is a notoriously fuzzy standard, but let’s put a figure on it and say that it corresponds to a 95% (0.95) probability of the factum probandum being true. To return to my murder trial, it would follow then that, in order to secure a conviction, the prosecution would need to introduce a body of evidence that (in its totality) raises the probability of my intentionally killing the victim to the 95% threshold. Contrariwise, it would also follow that if I could introduce any evidence that lowered the probability back down below the 95% threshold, then I would succeed in my defence. Thus, the probative value of the evidence varies depending on the context.

This is important because it feeds into the assessment of lie detection evidence. Reviewing the available literature, Schauer notes that fMRI lie detector tests have reported reliability rates that vary from 70-90%. This means they are better than chance at identifying deceptive individuals, but far from perfect. Unfortunately, in his discussion, Schauer doesn’t break down the data into false positives and false negatives. Consequently, I’m not unsure whether the 10-30% of failures covers truth-tellers who were falsely identified as liars or liars who were never spotted, or some combination of both. This could make a big difference to the legal utility of the evidence from the prosecutorial side in a criminal trial, but Schauer doesn’t look at the issue from their perspective.

Instead, Schauer looks at the issue from the perspective of the defence and notes that although a 70% reliability rate might not suffice to prove that someone is guilty, it might suffice to prove reasonable doubt. So, for instance, if I’m being tried for murder and I have an alibi which, following the administration of an fMRI lie detection test, is 70% likely to be true (Bayesian considerations to one side), it would be highly useful for the court to be made aware of this fact.

Breaking it down, the argument Schauer’s making looks something like this:


  • (1) In its present form(s), the accuracy rate of fMRI lie detection is somewhere between 70% and 90%. 
  • (2) In some probative contexts, a 70% likelihood that X is telling the truth/lying is highly probative. 
  • (3) Therefore, fMRI lie detection could be useful (even in its present form) in some probative contexts.


Thus we have the probative context argument. It should be pointed at that premise (2) can be defended with a number of examples. I used the criminal example since it’s possibly the most straightforward, but in civil trials the standard of proof is much lower (balance of probabilities) and hence the lie detector test could be highly probative in those contexts too.


2. Challenges to the Probative Context Argument
I have to say, Schauer’s basic point strikes me as being a good one. Nevertheless, there are some lingering concerns. Personally, I think the second premise needs to show some greater sophistication in its use of probabilities and accuracy rates. Thus, as mentioned previously, greater appreciation should be shown for rates of false positives and false negatives, not simply overall accuracy rates. A test that is 46% accurate might actually be highly probative, depending on whether the 54% of inaccuracies refers to false positives or false negatives. If the 54% refers solely to false negatives, then the test might actually be incredibly useful to the prosecution in a criminal trial. For in that case, the test would accurately identify guilty people to the exclusion of innocents. Thus, any concern about punishing the innocent would be allayed.

But this observation is a relatively minor one. The second premise of the argument could easily be reformulated and defended in such a way that the importance of false positives and false negatives is brought to the fore. A more pressing concern, and one that Schauer does actually address, arises in relation to the first premise. Critics will be keen to point out that the 70-90% accuracy rate is derived from experimental studies of the tests, not from real world applications. There are serious doubts as to the merits of extrapolating from such experimental studies to the real world. What might be 90% accurate in the laboratory setting, could be only 30% accurate in the field, or even less. We simply don’t know.

This is the ecological validity challenge. If it succeeds, it would undermine the probative context argument since that argument depends on us having some reasonable estimate of the accuracy of the test in question. If we have no such reasonable estimate — if the probabilities in question are, to put it bluntly, inscrutable — then Schauer’s argument won’t work. But are things really this bad?

Schauer thinks not. As he sees it, the ecological validity objection breaks down into two distinct parts. The first claims that we cannot extrapolate because the experimental subjects are not representative of the wider population. The second claims that the incentives under which people lie in an experimental setting are artificial, and quite distinct from the high stakes incentives in civil or criminal litigation.

Responding to the first claim, Schauer notes that this is a general problem with many kinds of evidence proffered for forensic use. For example, studies about the unreliability of eyewitness identification and memory are typically performed on undergraduate psychology students who may not be representative of the wider population. And because this is such a general problem, psychologists and other scientists have frequently sought to address it in their studies. They have done so by attracting more representative samples and trying to match real-world conditions more closely. Furthermore, they have tried to see whether results derived from the low-stakes unrepresentative sample tests hold up in the more high-stakes representative sample settings. Citing a slew of general reviews done on this topic, Schauer notes that the general trend seems to be that the results do hold-up. Although similar studies have not yet been done on fMRI lie detection, the trend may well remain the same unless there are particular difficulties with the extrapolability of fMRI results.

In relation to the second part of the objection, Schauer accepts the significant problems here. It is very difficult to artificially recreate the pressure to lie that might be felt in a real-world setting in the lab. But some fMRI researchers have tried to do this (Greene and Paxton, 2009) and their results are consistent with the premise underlying fMRI lie detectors. Future studies should address this in more depth and thus a more reliable picture of extrapolability can emerge.

A final related point emerges from the individual-population divide. Most fMRI studies, as well as most scientific studies, generate their statistical output by averaging over the population of experimental subjects. This leads to a classic problem in the legal context: how can this population-level data be probative in the individual case? After all, just because a test is 70% accurate across a population does not mean it is accurate for a particular individual in a particular case. So should the information be used at all?

Although this has been a surprisingly popular critique in legal circles, particularly when it comes to the use of epidemiological studies in tort law, it is flawed. As Schauer points out, the fact that for any random person plucked from the population, a particular test is accurate 7 times of 10 is probatively valuable given the right probative context. So this does not defeat the probative context argument.

The only problem with all this is that it might suggest a certain weakness in the argument. After all, given the right context, a test with an exceptionally low probability of being correct (say 5%) might be probatively relevant. Is this a reductio of the argument, or just a necessary truth about the nature of evidence and proof? I won’t answer that question here.


3. The Epistemic Progress Argument
On its own, the probative context argument has some value. But when coupled with the second argument, the argument from epistemic progress, it makes a good overall case for Schauer’s thesis. To explain the epistemic progress argument, I’m going to rely on some concepts from epistemic systems theory, which I’ve covered before on this blog.

To review, an epistemic system is any social system that (at least sometimes) generates judgments of truth or falsity. The legal trial is classic example since it generates judgments of truth or falsity concerning the factum probandum. Following Koppl’s schema, the epistemic efficiency of an epistemic system can be defined as follows:

Epistemic Efficiency: A measure of the likelihood of the system reaching a true judgment. Either 1 minus the error rate of the system; or the ratio of true judgments to total judgments.

And epistemic progress in this way:

Epistemic Progress: A system can be said to undergo epistemic progress whenever its epistemic efficiency is increased.

The basic idea is that epistemic progress is a good thing, and that any reform to the system that allows it to progress would be welcome. The key, however, is that epistemic progress is always assessed relative to the existing level of epistemic efficiency. Thus, if we wished to argue in favour of a particular reform, we would have to do so by directly referencing the current level of efficiency. This relativistic property of epistemic progress has one interesting effect: if the current level of epistemic efficiency is low, then a particular reform with an unimpressive level of overall accuracy, may nevertheless be warranted on the grounds that it still raises the efficiency of the system.

Unsurprisingly, Schauer argues that this is true in the case of fMRI lie detection. This gives him the following argument:


  • (4) If a particular reform to an epistemic system leads to epistemic progress, then it ought to be (all else being equal) welcomed. 
  • (5) The admissibility of fMRI lie detection evidence would lead to epistemic progress in the law.  
  • (6) Therefore, (all else being equal) fMRI lie detection evidence ought to be welcomed.


Schauer argues in favour of premise (5) by highlighting how existing methods of solving the false testimony problem are rather lacking. Historically, the administration of the religious oath was thought to incentivise truth-telling. In a culture in thrall to the fear of hell, this may have had some sway, but in its modern secular form the oath relies on the desire to be honest and the threat of perjury to do its work. Arguably, neither of these are particularly effective and certainly the oath has no known accuracy rate associated with it.

Robust cross examination is also often singled out as an excellent method for solving the false testimony problem. But this is highly suspect. As Schauer notes, cross examination may expose inconsistencies in certain cases, but is unlikely to do so in the case of the seasoned or practiced liar (movie depictions of the practice notwithstanding). In these cases we may be left with contradictory testimonies, which can be very difficult for a jury to assess. Furthermore, as with the oath, there are no known accuracy rates associated with cross-examination.

In light of these comparators, the admission of fMRI lie detection would seem to represent an improvement. Since it does have known accuracy rates, and since it can do something to break the deadlock between contradictory testimonies, it could lead to epistemic progress. Thus, the argument goes through.

Two caveats are in order here. First, in his defence of premise (5) Schauer may have missed out on other methods of solving the false testimony problem, ones which, although not currently used, would be more progressive than fMRI lie detection. This wouldn’t defeat the argument, but it might lessen its appeal since those alternatives would be the better bet. Second, the conclusion to the argument includes an “all else being equal”-clause. It might be possible for someone to argue that, in the case of fMRI evidence, all else is not equal. For example, they could argue that judges and juries are known to overvalue the results of fMRI studies, hence the admission of fMRI lie detection might do more harm than good. Schauer actually looks at this objection in the article, suggesting that it is ineffective, but I won’t cover that discussion here. I think this issue actually deserves a more detailed consideration, which I may (if the mood takes me) cover in a future post.


4. Conclusion
To sum up, Schauer’s thesis is that the admissibility of fMRI lie detection evidence cannot be determined solely on scientific grounds. He makes his case for this thesis with two arguments. The first — the probative context argument — claims that techniques with (scientifically) unimpressive accuracy rates might still be desirable in the legal setting. This is because the value of evidence varies with the probative context. The second — the epistemic progress argument — claims that even if fMRI evidence is not particularly reliable, its use in the law might nevertheless be desirable if it can raise the epistemic efficiency of the legal system. This, he argues, is something it could well do given that existing methods for solving the false testimony problem are rather weak.

Sunday, December 9, 2012

Third Anniversary



I almost forgot, but then I remembered. My first post on this blog was on Wednesday the 9th of December 2009. That makes this blog 3 years old today. I'm as shocked as anyone to find that my interest has been sustained over that time period.


Schauer on (fMRI) Lie Detection in the Law (Part One)



Regular readers of this blog will be aware of my interest in scientific evidence and law. As part of that interest, I have spent some time looking at the potential uses of neuroscience-based lie detection (or memory detection) tests in the law. In this post, I take up this interest again by looking at a recent article by Frederick Schauer entitled “Lie-Detection, Neuroscience and the Law of Evidence”.

The article is notable in two respects. First, it is an attempt by a leading scholar in the philosophy of law to weigh-in on an important issue in the study of “neuroscience and the law”. Second, unlike many who have written about this issue in the past, Schauer thinks that neuroscience-based lie detection could have an important role to play in the legal context, even in its present form. This sets him apart from others who, though noting the potential, typically deem such technologies “nascent” or “not yet ready” for legal use.

Schauer has pushed this case in the past, but the above-named article is his latest and most perspicuous defence of it. Over the next two posts, I want to clarify and formally reconstruct what I take to be the two major arguments in Schauer’s article. These are the probative context argument and the epistemic progress argument. To set these two arguments up, I first give some background on the use of evidence in the law, and on the history of the lie detector test. I then proceed to outline both arguments and address the key premises of each.

This post focuses on the background and history, leaving discussion of the arguments themselves ’til the second post.


1. Deception, Bias and Solomon’s Problem
Though it has oft been criticised as an inaccurate or incomplete characterisation, I believe it is true to say that most legal reasoning fits within a syllogistic pattern. That is to say, most legal cases revolve around the question of whether something like the following syllogism is true:


  • (1) If S did X, then legal consequence Y follows (legal rule)
  • (2) S did X (factum probandum
  • (3) Therefore, legal consequence Y follows (verdict/ruling).


Let’s take a very simple example. Suppose I am being tried for murdering my best friend. The governing legal rule in such a case would be (roughly): if a person (a) performs an act that causes the death of another person; and (b) they performed that act with intention to kill or cause grievous bodily harm, then they are guilty of murder. So if it could then be proved that I did perform such an act, and that I did so with the relevant intention, it would follow that I was guilty of murder.

Now, legal cases can often be more complex than this, with many chains and nests of syllogisms being linked together in one legal trial, but the basic pattern of reasoning remains the same. And it is this pattern, particularly the second premise in this pattern, that is important here. For it is this premise that states the factum probandum — the key legal fact that needs to be proved in each case — and this premise that reveals the allure of lie detection.

We can see this by considering the factum probandum in more depth. In order to prove this fact (or facts, as the case may be), the court relies on evidence. This evidence is usually presented to the court in the form of witness testimony. In other words, witnesses are put before the court to tell the court about what they saw or what they experienced or, exceptionally in the case of experts, to offer opinions about what might have happened. The problem is that, at least in common law systems, the system is adversarial. As a result, both sides present witnesses and these witnesses oftentimes contradict one another. Thus, it becomes difficult for the court to figure out where the truth really lies.

The contradiction is sometimes attributable to honest mistake, but other times is attributable to the strong incentive to mislead the court. After all, no-one likes to lose a legal case; everyone wants to win. This is classically illustrated in the biblical story of Solomon and the two women. One woman, who kills her child by rolling on top of it in her sleep, tries to claim the child of a second woman as her own. The second woman claims the child is really hers. They bring their case to Solomon who is asked to stand in judgment as to which woman should get to keep the child. But both witnesses contradict each other and, in the absence of further evidence, it is difficult to say who is telling the truth. This the problem of “false testimony”.

As we all know, Solomon solves this problem by changing the incentive structure of the case. He calls for the child to be divided in two and shared between the women, believing that this will incentivise the true mother to alter her testimony so as to benefit the lying mother. This she duly does and so she is awarded the child. This solution, though perhaps ingenious, is simply one of many. Another solution would be to have some credible, and reliable device or system for determining whether someone is telling the truth, one that does not rely on clever tricks such as Solomon’s alteration of the incentive structure. A lie detector test could do exactly this, hence its obvious allure in the legal system.


2. The Curious History of the Lie Detector Test
Despite the obvious allure of a reliable lie detector test, courts have typically been wary about admitting their results in legal cases, particularly in the U.S.. This trend was established in the very early days of the polygraph lie detector test. The polygraph test, invented in something like its modern form in 1921 by John Larson, records levels of physiological activity in the autonomic nervous system. In its classic form, it is based on the premise that elevated levels of activity in the autonomic system are reliable indicators of deception.

In a 1923 decision (Frye v. United States), the US Court of Appeals for the District of Columbia deemed that results from an early version of the polygraph test were inadmissible in a court of law. The case was important not just for this verdict, but for the fact that it set out the test for the admissibility of scientific evidence in US courts for the best part of 70 years: the general acceptance test (or Frye test, if you prefer). According to this test, scientific evidence, such as the result of a polygraph, was inadmissible if it was not generally accepted as reliable within the relevant scientific community.

As Schauer notes, the polygraph test has never really recovered from this early blow. Although the test has become more sophisticated, and although new methods of eliciting and identifying physiological signals that are thought to encode “deceptiveness” have come on stream, courts remain ambivalent, to say the least. This is despite its regular use in employment and non-legal settings. In law, doubts are still expressed about the test’s reliability and accuracy, and concern expressed about its potential to “unfairly prejudice” legal proceedings.

The doubts about the lie detector’s legal admissibility have become particularly significant in the past ten years or so. With the advent of neuro-imaging based techniques for lie-detection, coupled with a general fondness within academic and media circles for all things law and neuroscience-related, the debate about the admissibility of lie detection has caught the public eye once more. Several companies now offer fMRI-based lie detection services for use in legal trials, and some scientists argue strongly in favour of its forensic utility. But they have been met with fairly stiff opposition from other scientists and, indeed, from the courts. The most famous instance of this coming in the 2010 decision in United States v. Semrau which declared that the results of an fMRI-based test were inadmissible under the Daubert test for the admissibility of scientific evidence (this being the new, more sophisticated test that replaced the one put forward in Frye).

But is this scepticism warranted? Should courts be so reluctant to admit the results of an fMRI-based lie detection test? Or, indeed, any of the more modern variants of the test? Schauer suggests not. And we’ll see why he suggests this in part two.

Sunday, December 2, 2012

Can death be a fitting punishment?



This post is going to be about death, retribution, and the relationship between the two. It’s the last post I’m going to do on the philosophy of punishment for a while, but it deals with a significant claim that I’ve danced around in previous posts without addressing head on. What is that claim? It is the claim, common to at least some retributivist defenders of capital punishment, that death is the appropriate punishment for certain kinds of crime (most obviously: murder). But is this right? And how do we decide?

To set things up, we need to consider the basic thesis of the retributivist, which I’ll summarise as follows:

Retributivism: If a person engages in a culpable wrong of type X, then it is right and proper (perhaps obligatory) for them to be punished in a manner that befits that type of wrongdoing.

There are two key parts to this thesis. The first is the notion that punishment is intrinsically good. That is: good irrespective of its broader consequences. This is why it is right and proper (perhaps obligatory) to punish those who engage in culpable wrongdoing. The second is the notion that punishment must satisfy some fittingness-relationship. That is, to use a common formulation: the punishment must fit the crime (note “crime” is perhaps a little narrow since retributivism could cover all types of wrongdoing not just criminal wrongdoing, but that’s by-the-by since we will be focusing on criminal wrongs in the remainder of this post).

It is this second part of the retributivist thesis that I want to focus on today. For it is this part that motivates retributive justifications of the death penalty. Only if death is the fitting punishment for particular crimes will it be right to recruit retributivism in support of capital punishment. But there’s an immediate problem here: the notion of a “fitting punishment” looks to be somewhat vague. To justify their defence of the death penalty, retributivists will need provide some specification of the fittingness-relationship between a crime and a punishment that clearly implies that death is the appropriate punishment. Can they do this?

In the remainder of this post, I’ll try to answer this question by looking at two specifications of the fittingness-relationship. The first — called the “punishment-in-kind” version — is generally thought to be unacceptable even though it could justify the death penalty in particular cases. The second — called the “qualitative matching” version — looks to be more morally acceptable, but it faces problems of abstraction that weaken it’s ability to support the death penalty. And although I’ll refrain from drawing any broader implications from this, it does suggest that a retributivist might (a) struggle to justify the death penalty; or (b) consistently reject it.

[Source Note: This post is cobbled together from a variety of sources and personal reflections on the topic, but is heavily indebted to Jeremy Waldron’s article “Lex Talionis”]


1. The Punishment-in-kind Version
Here’s the first attempt to refine the fittingness-relationship:

Punishment-in-kind Principle: If a person engages in a culpable wrong of type X, then the fitting punishment is for a wrong of type X to be visited upon them.

The use of the word “type” is significant in this definition. In act theory, there is a distinction drawn between act types and act tokens. Roughly, an act type is a general classification or description, whereas an act token (or tokens) is a particular performance of an act type. To give an example, “buying a house” is a general act type, which is performed by the specific act tokens of me taking out a mortgage with my bank, and signing various contracts on a particular day. The same act type could be performed in myriad different ways by different people and at different times.

Some of the subtleties are irrelevant here. The important point is that the punishment-in-kind principle says that whenever a person performs a token of a particular general type of wrong, the same type of wrong must be performed to them. This seems to provide obvious support for the death penalty, at least on some occasions. If A stabs B to death, or if A shoots B, then A has performed the wrong of killing another. Therefore, by the punishment-in-kind principle, the fitting response is for them to be killed. Simple as that, right? The basic logic is illustrated in the diagram below. Two different murders are grouped within the same act type (which is clearly morally wrong) and therefore warrant the same response.



The problem is that the punishment-in-kind principle seems to lead to both practical and moral absurdities. Consider, for instance, the serial killer who has killed multiple people. Is the fitting response for them to be killed, revived and killed again for the relevant number of times? This would seem infeasible (though one could imagine a crude analogue in which the person has their heart stopped for a few minutes, before being revived and undergoing the process again). What about the killer who tortured his victim first? Should he be tortured and then killed? Or how about the rapist, should he be raped as punishment for his wrongdoing?

Superficially, the punishment-in-kind principle seems to warrant these responses. But this is surely absurd. A system of punishment that followed the principle to these extremes would seem unwarranted and downright inhumane. Nobody in their right mind could think it intrinsically good for the government to subject serial rapists to multiple rounds of rape.

So goes the standard objection to the punishment-in-kind principle. It leaves the retributivist with two options. Either bite the bullet and accept these troubling implications, while at least preserving their ability to justify the death penalty in some scenarios. Or seek an alternative specification of the fittingness-relationship that avoids these unpalatable consequences, hoping that the justification of the death penalty remains intact. We consider this possibility next.


2. The Qualitative Matching Version
Here’s the second attempt to specify the fittingness-relationship:

The Qualitative Matching Principle: If a person engages in a culpable wrong of type X, then the fitting punishment is for an act that qualitatively matches the wrong-making properties of X to be performed to them.

The tools of act theory can be used to flesh this principle out as well. Previously, we limited ourselves to one level of abstraction when seeking the fitting punishment. That is to say, we grouped particular act tokens — stabbing to death in one case, shooting in another — together into the general act type of killing. We then said that replicating this general act type would be the fitting response to those act tokens. This was to engage in one-level of abstraction in the search for fitting punishment (from the particular act tokens to the general act type).

But we could have gone further. We could have asked: what are the properties of the act type of killing that makes particular tokens of that act wrong? There are many different accounts of this. For instance, it could be that the general act type of killing is wrong because it exemplifies the even more general wrong of “permanently denying someone the capacity to consciously self-direct their life” or of “irreversibly terminating a person’s opportunity for future positive experiences”. If we looked for fitting punishments in these higher levels of abstraction, the goal of punishment would be to come up with some practice that qualitatively matches these kinds of wrongs.

We’ll return to these specific accounts of the wrongness of killing in a moment. But for now let’s consider whether this approach solves the problems with the punishment-in-kind principle. As we saw, the problem with that principle was that, if followed strictly, it seemed to warrant extremely harsh and inhumane forms of punishment. This was because the punishment had to replicate the general act type of the wrongdoer, as in the raping the rapist counterexample. Arguably, this problem was caused by the fact that it limited itself to the first level of abstraction. If we jumped to a higher level of abstraction, we wouldn’t have to resort to inhumane punishments of this sort. Thus, the wrongness of rape could be that it exemplifies the more general wrongs of “violating autonomy” or “domination and exploitation”. It is possible to replicate those wrongs in a punitive act without actually trying to rape the wrongdoer.

I don’t want to belabour this example too much because it merely meant to illustrate how the qualitative matching principle might work. The key question here is whether the principle supports the death penalty. The answer is far from clear. If we think the wrong-making property of killing is that it permanently denies someone the capacity for conscious self-direction, then punishments short of death may be fitting. For instance, inducing coma might do the trick. Likewise, if we think the wrong-making property of killing is that it irreversibly terminates the opportunity for future positive experiences, then life in prison, under harsh conditions, without the possibility of parole, might do the trick.

The point is that once we jump to the higher levels of abstraction, the wrongness we are trying to match will be capable of being exemplified in many different kinds of act tokens. That is the usual effect abstraction: it groups more and more discrete phenomena together under a general category or label. As a result, the justifiability of the death penalty becomes less immediate. The defender will need to show how alternative types of punishment, which seem to qualitatively match the wrongness of the crime, actually fail to do so and that only death lives up to the demands of the principle. That’s a difficult task since it requires the exhaustive consideration of all the possible alternatives.


3. Conclusion
To sum up, the retributivist defender of the death penalty bases their defence on the claim that death is the fitting punishment for certain kinds of crime (most obviously: murder). The problem with this claim is that the notion of a “fitting punishment” is vague. One way to specify it would be to appeal to the punishment-in-kind principle of fittingness. But, as we saw, this seems to lead to absurd and unwelcome conclusions. Such as: the fitting punishment for a torture-murderer is for them to be tortured and then murdered, or the fitting punishment for the serial rapist is for them to repeatedly raped.

One suggested diagnosis of the problem with the punishment-in-kind principle is that it limits itself to one level of abstraction in the search for the fitting punishment. If it were reformulated so as to allow a search through higher levels of abstraction, the absurd and unwelcome forms of punishment could be avoided. But once we do this, as we did with the qualitative matching principle, it becomes much less clear that death is uniquely warranted by retributivism. Thus, the earlier claim seems to hold: it is possible for a retributivist to (a) struggle to justify the death penalty; or (b) consistently reject it.

Tuesday, November 20, 2012

The Reversal Test and Status Quo Bias



Changing policies can often seem arduous, and undesirable, even when it might be for the best (ethically speaking). As a result, reluctance to change can often creep into many organisations. I’m sure we’ve all encountered it. This reluctance is compounded by two other facts. The first is that we are usually deeply uncertain about the long-term consequences of any proposed reforms to the systems in which we operate. Consequently, when we reason about such things, we tend to fall back (at least in part) on our intuitive judgments about what seems right and wrong. The second fact is that, as numerous studies in cognitive psychology bear out, humans seem to be intuitively biased in favour of the status quo. So when our uncertainty forces us to rely on our intuitions, reluctance change is the natural result.

In an article written several years back, Nick Bostrom and Toby Ord argue that this bias to the status quo is a major problem in applied ethical decision-making. In order to be rational ethical decision-makers we ought to systematically check ourselves against the possibility that our aversion to a particular policy is driven by the bias toward the status quo. To do this effectively, they propose the introduction of something called the Reversal Test. In this post, I want to explain what this test is and how it works. As we shall see, there are really two tests, and they each have slightly different effects.

Before I begin, I should acknowledge that some may doubt the existence of a systematic bias toward the status quo. To them, much of what follows may seem unjustified. But the evidence for the status quo bias looks to be abundant and robust. Bostrom and Ord discuss this evidence in their article, and presentations of it can also be found in Kahneman’s work Thinking Fast and Slow. Although I am happy to entertain doubts about this evidence, I shan’t discuss it here. Instead, I’ll skip directly to the tests themselves, since that’s where my interest lies.


1. The Reversal Test
The reversal test is, in essence, a heuristic or rule of thumb that counteracts the effects of the status quo bias. The test can be stated like this:

The Reversal Test: When a proposed change to a certain parameter (in a certain direction) is thought to have bad overall consequences, consider a change to the same parameter in the opposite direction. If this is also thought to have bad overall consequences, then the onus is on those who believe this to explain why any changes to the parameter are deemed to be bad. If they are unable to do so, we have reason to suspect they are suffering from the status quo bias.


Thus, to give an overly-simplified example, suppose we are being asked to consider a proposed increase in the speed limit (from, say, 60mph to 70mph) and most people seem to think this would be bad. Then, we ask them whether a reduction in the speed limit (from 60mph to 50mph) would also be bad. If they think so, we ask them to justify their belief that 60mph is the optimum speed limit. If they cannot, we have reason to suspect they are biased toward the existing status quo.

On the face of it, this is pretty banal. We are just asking people to justify their beliefs which is surely what they should be doing anyway. Nevertheless, its practical effect could be significant. This is because the Reversal Test performs one crucial function: it shifts the burden of proof. Typically, we think that the burden of proof is on those proposing change. But assuming they can offer some reason for the change, and are nevertheless resisted, the Reversal Test has the neat effect of shifting the burden of proof onto the resisters. They have to explain why the current state of affairs represents a local (or absolute) optimum within the possible space of parameter values.


2. The Double Reversal Test
Of course, the burden of proof could be met. In particular, opponents of the policy could point to risks inherent in the proposed changes, or to transition costs that outweigh the value of the proposed changes. But there are problems with these kinds of responses too. Namely: humans are not particularly good at estimating risks and, due in part to status quo bias, they tend to overestimate the actual costs associated with proposed changes, focusing too much on short-term transition costs and not enough on potential long-term benefits.

So Bostrom and Ord propose an extended version of the test, which they call the Double Reversal Test:

Double Reversal Test: Suppose there is resistance to changing the value of a parameter in any direction. Now imagine that some natural event threatens to change the value in one direction. Would it be a good thing to counterbalance the effect of that natural event with something that maintains the status quo? If so, then ask whether, assuming that the natural event reverses itself at a later point in time, it would also be a good idea to reverse the counterbalance so as to maintain the original value of the parameter? If no one thinks so, then current opposition to the policy is likely to stem from the status quo bias.

The basic idea behind this test is illustrated in the diagrams below.




Although the diagrams help, this version of the test is difficult to follow in the abstract. Fortunately, Bostrom and Ord give quite a nice example of what it really means. Their example concerns resistance to cognitive enhancement technologies, which is, in fact, their focus throughout the paper. They think that current opposition is driven largely by the status quo bias, and not by any coherent moral principle. So they ask us to imagine the following scenario.

A hazardous chemical has entered the municipal water supply. Try as we might, there is no way to remove it, and there is no alternative water source. The chemical has the disastrous effect of impairing everybody’s cognitive function. Fortunately, there is a solution. Scientists have developed somatic gene therapy which will permanently increase the cognitive function of the population just enough to offset the impairment caused by the chemical. Everyone breathes a sigh of relief; the current level of cognitive capacity is maintained. But, at a later time, the chemical begins to vanish from the water. If we do nothing, cognitive capacity will be increased over its original level. So should we do something to reverse the effect of the somatic gene therapy? If not, then it’s likely that current opposition to cognitive enhancement stems more from status quo bias than from any coherent moral concerns.

The Double Reversal test is because it helps to disentangle two distinct conceptions of the status quo:

The Average Value Conception: In which the status quo is viewed as the current average value of the parameter in question.
The Default Position Conception: In which the status quo is viewed as the value of the parameter if no actions are taken.

Allegiance to the current set of average values might be ethically justified, and if we are willing to intervene to counterbalance the natural event, then perhaps we have some principled reason to think the current average is optimal. But if we don’t think it is necessary to counterbalance the original policy after the natural event reverses itself, then we are switching to a default position conception of the status quo. Switching in this manner suggests our attachment to the current set of values is unprincipled. After all, if we are willing to take the risk and incur the transition costs to counterbalance the natural event, but unwilling to incur additional costs to counterbalance the subsequent reversal of the natural event, then what is current opposition really based on?

In sum then, Bostrom and Ord’s reversal tests are useful heuristics to employ in ethical policy-making. The basic Reversal Test is useful because it shifts the burden of proof onto those who defend the status quo, and the Double Reversal Test is useful because it allows us to see more clearly whether the current attachment to the status quo is principled or not.

Saturday, November 17, 2012

Is the Death Penalty Irrevocable? (Part Two)



(Part One)

This is the second part in a brief series of posts looking at Benjamin Yost’s discussion of the Irrevocability Argument against capital punishment. As explained in part one, the Irrevocability Argument claims that the death penalty is a morally illegitimate system of punishment because it is not substantially revocable. That is to say, unlike other forms of punishment which can be corrected if they are wrongfully imposed, the errors of wrongful execution cannot be corrected. Once a person is dead, they’re dead. You cannot make it up to them.

In part one, we looked at Michael Davis’s objection to the Irrevocability Argument. According to Davis, the death penalty is substantially revocable because it is possible to benefit a person after they die. This can be supported by direct appeal to the Pitcher-Feinberg theory of posthumous harms. This theory holds that a person is benefitted if their interests are satisfied or fulfilled. And since a person’s interests can outlast their physical lives, it follows that they can be benefitted after they die. Davis merely adds to this the claim that the posthumous benefit can be sufficient to outweigh or counterbalance the harm done by their execution. Thus, we get the following argument:


  • (4) In order for a system of punishment to be substantially revocable, it must be possible to compensate people after their punishment such that the wrong done to them by the punishment is outweighed or counterbalanced by the compensation (Substantial Revocability principle) 
  • (5) It is possible to compensate people (greatly) after they die. (Pitcher-Feinberg theory) 
  • (6) Therefore, the death penalty is (in principle) substantially revocable.


Clever and all as this argument is, it is open to at least two criticisms. The first targets premise (4) and argues that substantial revocability cannot be reduced to compensation. The second targets premise (5) and argues that the Pitcher-Feinberg theory is flawed. We’ll look at both today, in reverse order. As we shall see, in his analysis, Yost thinks the first criticism is the better bet, but it’s worth exploring the second one anyway.


1. Problems with the Pitcher-Feinberg Theory
The Pitcher-Feinberg theory analyses harm and benefit in terms of the interests and preferences of the person. A thwarted interest is a harm, and a satisfied interest is a benefit. Using carefully crafted thought experiments, like the “Mortal Metaphysician” example discussed in part one, proponents of this theory show that posthumous harm/benefit is plausible. But despite the prima facie plausibility of these thought experiments, two critical questions can be asked:

(A)The Subject Question: Who exactly is being harmed or benefitted in this situation?
(B)The Causation Question: If a person can be harmed or benefitted after they die, does this not require some spooky backwards causation?

The questions are connected, in that answering the first in a particular way concedes some ground to the premise of the second question. If one claims, as Pitcher does, that the only plausible candidate for a subject of harm/benefit is the ante-mortem person (since the post-mortem person is just dust in the grave), then one naturally encounters the charge of backwards causation. But how serious a charge is this?

Thought experiments and analogies ride to rescue here. Consider the following scenario: Shortly after Barack Obama’s exit from the presidency in 2016 (or, rather, Jan. 2017) the world is struck by a gigantic meteor that wipes out all of human civilisation. This has the curious effect of making Barack Obama the penultimate president of the United States. In other words, it adds a property to the pre-2017 presidency that wasn’t there before. But this doesn’t require any spooky backwards causation. Why couldn’t posthumous harms be the same? In other words, why couldn’t attaching the properties of harm and benefit to the ante-mortem person be like attaching the property of “penultimacy” to the president?

Here, we get into a war of intuitions and thought experiments. Critics like James Stacey Taylor argue that the property of penultimacy is unlike the property of harm/benefit in that it is a sequential property and they are not. This may indeed be true, but it’s not at all clear that this is a relevant disanalogy, i.e. one that undermines the original claim. What may be going on is that there is a battle being waged between two different theories of harm and benefit: the interest-based account of Pitcher and Feinberg, on the one hand, and the experiential or Epicurean account, on the other.

Classically, a person is defined as a continuing subject of experiences, or an overlapping set of psychological states. According to the experiential account of harm/benefit, a person can only be harmed/benefitted if there is some change in those experiences. So, for example, according to this account, my wife’s death harms me if (and only if) I become aware of it. This conception does not allow for posthumous harm/benefit. After all, once the person is dead, the continuing subject of experiences ceases to exist, and so there can be no further changes to what they experience. Another way of putting it is that this view requires some intrinsic change in the person before there can be harm or benefit. Contrariwise, the Pitcher-Feinberg theory allows for extrinsic changes to harm or benefit the person.

At this point, one is reduced to a battle between two opposing views, with little common ground between them. Defenders of the Pitcher-Feinberg theory will cling to their thought experiments and the seeming plausibility of their view, while defenders of the experiential account will cling to the intuitive good sense of theirs. Shelley Kagan tries to broker peace between the two sides by distinguishing between two potential loci of harm/benefit: (i) lives; and (ii) persons.

According to Kagan, every person has a life (a biography or set of interests) that it is possible to harm (or benefit) extrinsically; while at the same time the person themselves is indeed a set of experiences that can only be harmed (or benefitted) intrinsically. Thus, for Kagan, the Pitcher-Feinberg thought experiments provide support for the view that lives can be posthumously harmed/benefitted, whereas it remains true that persons cannot.

Where does this leave us? Well, if one accepts Kagan’s distinction between lives and persons, Davis’s challenge to the Irrevocability Argument is still viable. The premises need to be rewritten to acknowledge the distinction, but one could nevertheless argue that a person’s life can be compensated after they die (as the discussion in part one suggested). But there are still difficulties, in particular one might argue that even if lives can be benefitted, the notion of compensation is distinct. Only persons can be compensated. That would be an interesting criticism, but Yost avoids it because he thinks there is a bigger problem with the reduction of substantial revocability to compensation. We close by looking at that problem.


2. Substantial Revocability, Compensation and Control
Premise (4) states that a punishment can be substantially revoked if the harm done by the punishment is outweighed or counterbalanced by the benefit of the compensation. That sounds plausible enough until you probe a little deeper. When you do, you’ll discover, as Yost argues, that revocability cannot be reduced to compensation.

This discovery is prompted by the fact that, if revocability were reducible to compensation, it would be possible to compensate someone who was imprisoned, but still keep them in jail. Consider once more the case of the single father, who’s one wish in life is to make sure his daughter has a better life than him. Suppose he is wrongfully imprisoned, and the error is later discovered. He initially looks forward to his release, but is told that this would send a bad signal to other would-be criminals. The state is trying to look tough on crime, and their team of psychological advisors have informed them that releasing a prisoner, even a wrongfully imprisoned one, would damage the deterrent effect of incarceration. So the man is offered a deal: if he stays in jail, the state will provide a top-class education for his daughter, and make sure she secures a high-powered position within the civil service. The man readily agrees.

In this hypothetical scenario, the compensation paid to the man’s daughter would clearly seem to outweigh (or counterbalance) the harm done to him by the incarceration. But he nevertheless continues to be punished. Surely this is absurd? Surely one cannot substantially revoke a punishment, while leaving it in place? Thus, we seem to have a reductio of premise (4):


  • (7) It would be absurd if a punishment could be substantially revoked while nevertheless being left in place. 
  • (8) If substantial revocation is reduced to compensation, it would be possible to substantially revoke a punishment while nevertheless leaving it in place. 
  • (9) Therefore, the reduction of substantial revocation to compensation is absurd.


Thus, premise (4) is rebutted. One could leave it there since Davis’s argument depended on that premise, but one would be unwise to do so. If one did, it would remain open to someone to invent a new account of substantial revocation which corrected for the flaw in Davis’s account. Yost avoids this by providing an additional argument to the effect that an essential element of substantial revocation would never be possible in the case of the death penalty.

What is this essential element? It is the restoration of control or autonomy. As Yost sees it, modern liberal democracies are founded on respect for the moral and practical autonomy of their citizens. The foundational principle of most liberal theories (e.g those propounded by Hobbes, Rawls or Gaus) is that people are moral equals. Which is to say, no one can claim coercive moral authority over another without good moral reason. So one of the legitimacy conditions for a state is that it respect the moral autonomy of its citizens: if it is going to restrict or coerce them in some way, it better have a damn good reason for doing so.

Punishment is an exercise of coercive moral authority: in punishing a person, the state deliberately harms and coerces them in order to serve some moral end. As such, they violate the moral autonomy of the person being punished. If the state is supported by good moral reasons when doing so, then maybe that’s okay. But if the state makes an error, the harm done by the violation of moral autonomy must be corrected. Yost argues that the only way to do this is by restoring control to the person punished. (Note: his account of control is metaphysically modest, assuming only that compatibilist control is possible).

Now comes the clincher: you cannot restore autonomy or control to a dead person. They are dead. So even if you could compensate them by satisfying their interests, you could never substantially revoke their punishment in the manner required in a liberal democracy. In other words:


  • (10) In order for a system of punishment to be substantially revocable, it must (at a minimum) be possible to restore control to the person who was punished. 
  • (11) One cannot restore control to a dead person. 
  • (12) Therefore, the death penalty is not substantially revocable.


To be clear, Yost is not saying that the restoration of control is all that is required for substantial revocation, compensation could well be part of the picture. What he is saying is that restoration of control is necessary for substantial revocation, and that is enough to defeat Davis’s argument.




3. Conclusion
To sum up, the Irrevocability Argument is a powerful, conceptual argument against the death penalty. If one accepts that, in order to be morally legitimate, a system of punishment must include some capacity for error correction or revocation, then the fact that the death penalty is irrevocable counts against it. But, as we have seen, it is possible to challenge this argument by claiming that, contrary to what one might think, the death penalty is substantially revocable.

In this series of posts, we have seen how it is possible to defend this notion by reducing revocation to compensation, and arguing that one can harm or benefit a person after they die. However, we have also seen that the theory of posthumous harm and benefit is open to criticism, and that the reduction of revocation to compensation is not satisfactory. If these critiques are right, then the Irrevocability Argument is left standing. A useful weapon in the arsenal of the abolitionist.