Introduction
Opaque models now influence a staggering number of consequential decisions. However, even experts struggle to understand exactly how and why these models make the decisions that they do. As such, many argue that no explanation is available for decisions made or substantially influenced by opaque models.1
Some find the purported inexplicability of these decisions to be problematic. Vredenburgh (2022), for one, contends that decision-subjects have a right to an explanation of certain institutional decisions. She argues that this right protects their interest in informed self-advocacy, and when opaque models are used as input to these decisions, the right is threatened. Accordingly, Vredenburgh thinks that we must render such decisions explicable. We can do so using a variety of methods in explainable AI (XAI), such as fitting simpler models over opaque ones to approximate these models’ decision-making process.2
Vredenburgh also recommends that institutions provide decision-subjects with access to free expert advocates. Interpreting the outputs of XAI methods often requires some degree of statistical proficiency, which cannot be reasonably expected of decision-subjects. Accordingly, institutions should provide them with access to free experts who can use their knowledge of XAI methods and the relevant institutions to interpret the outputs of these methods. The experts, as Vredenburgh envisions, then use this information to advocate for decision-subjects and to detect any potential malfeasance on the part of decision-makers.
Grote and Paulo (2025) challenge this view. They argue that informed self-advocacy cannot be reliably promoted by the explanations that Vredenburgh seeks, even when institutional decision-makers provide decision-subjects with access to expert advocates. These explanations will either be too coarse-grained to enable meaningful self-advocacy or too fine-grained to be intelligible to decision-subjects. Moreover, they argue that expert advocates will not help matters as the proposal is likely to be both ineffective and intolerably costly.
Grote and Paulo note that a counterfactual approach—which informs individuals how a model’s inputs would have to change to yield a different output—could dissolve this dilemma. However, they contend that counterfactual methods must satisfy very specific conditions to do so and accordingly remain skeptical of the approach.
I think Grote and Paulo are, at first glance, justified in their skepticism of counterfactual methods. Karimi et al. (2021), for instance, argue that the standard approach to counterfactual explanation treats input features as causally independent, yielding unrealistic counterfactuals that do a poor job of informing decision-subjects what they can do to achieve their desired outcome. But informed self-advocacy requires, inter alia, that decision-subjects know what they can change to achieve a more favorable outcome.3 So, the standard approach seemingly fails to enable one aspect of the self-advocacy that Vredenburgh’s right serves to protect.
I will argue that this skepticism can be overcome. By employing a different approach to counterfactual explanations—one using variational autoencoders (VAEs)—decision-makers can provide explanations that satisfy an essential part of Vredenburgh’s right. Because VAEs learn realistic distributions of input data, they generate counterfactuals that represent profiles real people could actually have, avoiding the unrealistic counterfactuals that Karimi et al. warn against. Counterfactuals generated using the VAE approach therefore better promote one dimension of informed self-advocacy and thus help institutional decision-makers fulfill their explanatory obligations to decision-subjects.
Right to an Explanation
I will begin by outlining Vredenburgh’s account. On her view, individuals have a right to an explanation of adverse, institutional decisions that seriously and unavoidably impact their lives.4 Worryingly, institutions increasingly use opaque models to provide input to these adverse decisions. In such cases, individuals’ right to an explanation is plausibly infringed since, on Vredenburgh’s view, decisions made by these models are inexplicable.
The most significant move in Vredenburgh’s argument is establishing the right to an explanation. She does so within a Scanlonian framework. A right is justified when it protects a weighty and widespread interest, does so better than alternatives, and imposes costs that no one can reasonably reject. The interest at stake is what she calls informed self-advocacy—a cluster of abilities to “represent one’s interests and values to decision-makers and to further those interests within an institution.”5
Informed self-advocacy, as Vredenburgh conceives of it, has three components. The first, representation, concerns having one’s interests taken into account when institutional rules are put in place. For example, we have an interest in influencing the criteria an employer uses to evaluate candidates. The second, accountability, is the ability to hold decision-makers responsible when they make unfair decisions, such as when a loan applicant challenges a denial by alleging that discriminatory criteria were used in the decision-making process. The third, agency, is the ability to navigate institutional rules in order to achieve one’s goals—e.g. we have an interest in knowing what criteria must be satisfied in order to secure a loan, and what we can do to satisfy those criteria.
For our purposes, I will focus only on agency. Vredenburgh argues that this dimension of self-advocacy is best enabled by causal explanations that cite the relevant institutional rules alongside population-level causal generalizations about how different types of people fare in social systems.6 For example, a bank might explain that loan decisions are based primarily on debt-to-income ratio, and that applicants who reduce their debt-to-income ratio below a certain threshold are more likely to be approved. This explanation tells decision-subjects both what the relevant rules are and what they can do to satisfy them.
Grote and Paulo (2025) argue that the causal explanations required by Vredenburgh’s account cannot reliably be provided. If the causal explanations are to be intelligible to most decision-subjects, then they will be too coarse-grained to guide meaningful interventions. After all, population-level generalizations, by definition, refer to properties of larger social groups, but agency requires that a particular decision-subject understand what she can do to achieve a different outcome. Accordingly, these generalizations must be fine-grained enough to inform individual decision-subjects if they are to promote self-advocacy.7
However, producing fine-grained, population-level causal generalizations seems intolerably costly. Vredenburgh suggests that we ought to obtain these generalizations from the social sciences—i.e. through experimental or observational studies. But these studies, if they are to yield fine-grained causal generalizations, must split participants into a myriad of narrow subgroups, which places unrealistic demands on researchers.8
One might instead take a different approach to satisfying the agential component of the right to explanation. Grote and Paulo note that model-to-model explanations—those which expose the internal causal structure of a model rather than the real-world relationship between its input variables—could also promote this dimension of self-advocacy.9 But current XAI methods that exemplify this approach, such as linear approximation, are not sophisticated enough to produce genuine causal explanations.10
Moving beyond the state-of-the-art, however, it is plausible that methods of generating such explanations will eventually exist. Even if this is true, Grote and Paulo argue that the causal explanations generated using XAI techniques will be too complex for most decision-subjects to interpret. Opaque models are, at their core, statistical functions, so any satisfactory explanation that approximates the decision-making procedure of these models will require that individuals possess a level of statistical knowledge higher than can reasonably be expected.11
In response to this problem, Vredenburgh suggests that decision-subjects be provided with free expert advocates to help interpret model-to-model explanations. The advocates would function as explanation aides—much like caseworkers in welfare systems who help individuals understand why they received an adverse outcome and what they can do to contest it or alter their future behavior. These experts would have experience with how a particular institution works, and after supervising many individual cases, they would be well positioned to identify patterns of mistakes or unfairness across decisions in a given institution. They could then analyze the explanations for indications of wrongdoing and accordingly guide decision-subjects in light of their analysis.12
Grote and Paulo argue that, even with the help of expert advocates, Vredenburgh’s right cannot reliably be fulfilled by this approach to model-to-model explanations. They contend that the financial costs of retaining expert advocates will scale dramatically. Since the right to an explanation applies to all high-stakes algorithmic decisions, the demand for qualified advocates will soon outstrip the supply. After all, it seems that institutions are using opaque models to influence an increasing number of consequential decisions. This raises the question of who will bear the costs of these experts. If the costs fall on the institution, the right to an explanation risks becoming intolerably expensive.13
One might argue that this objection can be overcome. Perhaps there are more people with the relevant expertise than Grote and Paulo assume. Moreover, it is reasonable to think that the labor and capital that institutions save by automating decisions justifies the costs of expert advocates.
Even if Grote and Paulo’s objection is not decisive, there is further reason to doubt that we have a right to the explanations that Vredenburgh proposes. Recall that her argument operates within a Scanlonian framework: we have a right to a particular form of explanation only if it protects our interest in self-advocacy better than alternatives. Grote and Paulo themselves suggest that counterfactual explanations may promote agency while avoiding the problems they raise for other forms of explanation. And if counterfactual explanations do not face similarly weighty objections, then they plausibly better protect our agential interests. Accordingly, it would be counterfactual explanations to which decision-subjects have a right.
It should be noted that Grote and Paulo remain skeptical of counterfactual approaches. They contend that such approaches must satisfy very specific conditions to fulfill the right to explanation, and that current methods of generating counterfactual explanations face numerous problems.14 Even so, I posit that we should not be too hasty in abandoning counterfactual approaches.
In what follows, I pick up where Grote and Paulo leave off. I begin by outlining the standard approach to counterfactual explanation and show how it ostensibly serves the aims of informed self-advocacy. I then present an objection raised by Karimi et al., which demonstrates that the standard approach can generate unrealistic counterfactuals that are of little use to decision-subjects when seeking to achieve their desired outcome. Finally, I argue that an alternative approach—one that employs variational autoencoders—can overcome this objection and thus better promote agency than alternative forms of explanation.
The Standard Approach
I will now give an overview of the standard approach to generating counterfactual explanations for decisions made by opaque models. Wachter et al. (2018) offer the most comprehensive account. Their method, dubbed “the standard approach,” is most easily illustrated by way of example.
Imagine Mary applies for a loan at a bank that uses an opaque model to assess the loan-worthiness of applicants. The model takes variables such as salary, savings, income-to-debt ratio, and age as input. It then returns some output that the bank uses to deny Mary’s loan. Upset about her failure to secure a loan, she demands an explanation. The bank accordingly supplies one by providing counterfactuals representing the smallest change to Mary’s application that would have resulted in approval rather than denial. The counterfactuals may indicate that, holding all other variables constant, Mary would have been accepted if her income had been $2,000 higher. Similar counterfactuals could be generated for every input in the features space, and these counterfactuals jointly constitute an explanation because they inform Mary what she would have to change about her application in order to get approved for the loan.
The bank generates these counterfactuals by using an optimization algorithm that measures distances between actual and perturbed input variables, identifying the smallest changes to Mary’s actual profile that would yield a different output. Accordingly, the standard approach can be viewed in much the same light as adversarial perturbation: we perturb features of a particular input by small amounts until a corresponding change in output occurs.
From here, it is easy to see how counterfactual explanations of this kind could promote the agency component of Vredenburgh’s right to an explanation. By presenting a decision-subject with a minimally different version of their profile that would have received a favorable outcome, the explanation tells them what they would need to change in order to achieve such an outcome. And since a central part of agency, as Vredenburgh conceives of it, consists in knowing what one must do to further their interests within an institution, counterfactual explanations could satisfy this condition of informed self-advocacy.15
However, some have objected that the standard approach cannot effectively promote Vredenburgh’s conception of agency. Because it does not account for how realistic certain combinations of features are or how some features may depend on one another, the approach can yield unrealistic counterfactuals that mislead decision-subjects about what they would need to change to receive a different outcome. I now consider one such objection.
An Objection to the Standard Approach
Karimi et al. argue that the standard approach to generating counterfactuals does not provide sufficient guidance for decision-subjects to alter their circumstances in light of an adverse outcome. Such an explanation could, for example, mislead a rejected loan applicant about the optimal path to approval.
The standard approach can be misleading in this way because it does not effectively model the causal relations between features. Recall that the approach generates counterfactuals by perturbing input variables just enough for the output to change. Because these perturbations treat each variable as independent—e.g. a change in income would not reflect a likely corresponding change in savings—the resulting counterfactuals might represent combinations of features that no real person could have. And when a counterfactual is unrealistic in this way, it becomes unclear, from the decision-subject’s point of view, what she can do to flip the model’s output.
Karimi et al. provide an example to illustrate this problem. Imagine that Mary earns $75,000 annually, has $25,000 in savings, and is denied a loan. The opaque model used to evaluate loan applicants bases its decision on a combination of salary and savings. Here, the standard approach identifies the nearest counterfactual in which Mary gets accepted: she must increase her savings to $35,000 or her salary to $120,000. Because increasing her savings by $10,000 seems like the easier choice, Mary believes that she should pursue the former.16
However, suppose that in the real world, savings are causally related to salary: having a higher salary generally causes people to save more. So if Mary increased her salary to $85,000, her savings would likely increase as a consequence, and this combined change would suffice to flip the model’s decision and would potentially require less effort than saving an additional $10,000 at her current salary.
The standard approach fails to identify this less costly path because it treats each feature as independent. It does not represent the causal relationships between input variables, and in this case, savings and income are causally related. As a result, it can recommend changes that are more costly than necessary. As Karimi et al. argue, the standard approach tells individuals “where they need to go, but not how to get there.”17
One might object that Karimi et al.’s argument shows only that the standard approach provides sub-optimal guidance—not that it fails to promote agency altogether. After all, agency does not require that decision-subjects receive the best possible advice. Plausibly, it requires only that they have some idea of what they can do to achieve a more favorable outcome.18 If the standard approach tells Mary to increase her savings by $10,000, and she can feasibly do so, then perhaps the agential component of Vredenburgh’s right is fulfilled—even if there are more efficient ways for her to achieve the desired outcome.
This objection fails for at least two reasons. First, the standard approach does not merely produce sub-optimal guidance; it can lead decision-subjects to take infeasible actions. Consider the following case.
Suppose an opaque model evaluates loan applicants using age and length of credit history. Sarah, a 22-year-old applicant, is denied a loan. The standard approach identifies the nearest approved counterfactual: a profile identical to Sarah’s except with a credit history of five years instead of four. But since credit accounts cannot be opened before the age of 18, the longest credit history a 22-year-old could possibly have is four years.
Because the standard approach treats each feature as independent, it does not register that age constrains the possible length of one’s credit history. Accordingly, the resulting counterfactual does not describe a profile that is merely costly to achieve but rather one that is entirely infeasible. And surely, the agential component of the right to explanation cannot be fulfilled by an explanation that provides infeasible guidance. So, there are at least some cases for which the standard approach cannot fulfill the agential component of the right to explanation.
Second, setting aside infeasibility, the objection fails nevertheless. Recall again that the right to explanation, as Vredenburgh conceives of it, must protect decision-subjects’ interest in self-advocacy better than alternatives. So, regarding agency, decision-subjects only have a right to the form of explanation that best promotes their ability to navigate institutions in accordance with their interests. As I will now argue, VAE-based counterfactuals avoid Karimi et al.’s objection and thus better promote agency than the standard approach. Moreover, because the approach is still counterfactual, it sidesteps Grote and Paulo’s objections to causal generalizations. If I am correct, then, we have a right to counterfactual explanations generated using the VAE approach—as opposed to those generated by the standard approach or the explanations that Vredenburgh proposes—because they better protect the agential interests of decision-subjects.19
The VAE Approach
I will now illustrate how variational autoencoders can be used to produce counterfactual explanations that avoid the problems raised by Karimi et al. It will first be helpful to understand how VAEs work. Accordingly I will begin by describing ordinary autoencoders, the models on which VAEs are based.
An autoencoder is a model that learns to compress and reconstruct data. It consists of two parts: an encoder and a decoder. The encoder takes as input a vector that represents some real-world data point—say, a loan applicant’s profile, with features like income, debt, and credit history—and compresses it into a point in a lower-dimensional space called latent space.
The decoder then reverses this process. It takes a point in the latent space and converts it back into feature space, thereby reconstructing a full applicant profile. The model is trained so that the reconstructed output matches the original input as closely as possible, which forces the encoder to preserve whatever structure in the data is most important for accurate reconstruction.
To illustrate, consider an autoencoder trained on images of cats. The encoder compresses a high-resolution image, composed of thousands of pixels, into a much smaller list of numbers in latent space. If the model is well trained, these numbers might represent features important to classifying cats, such as fur color, ear shape, etc. The decoder would then reconstruct a full image from these latent variables. So, essentially, the latent space serves as an efficient summary of what matters most about the input data.
Importantly, however, an ordinary autoencoder imposes no constraints on the latent space. The encoder can place inputs wherever it likes, with each vector in input space being mapped to a single point in latent space. As a result, the latent space often contains gaps—regions that do not correspond to any real data point. If we were to pick a point in one of these gaps and decode it, the output might be nonsensical—such as a profile with an impossible combination of features. This makes ordinary autoencoders poorly suited for generating counterfactuals, since the search for a nearby accepted profile might land in one of these gaps and thus produce an unrealistic output.
VAEs address this problem by adding a constraint to the bottleneck layer. Rather than mapping each input to a single point in the latent space, the VAE’s encoder outputs two values for each latent variable: a mean and a standard deviation. Together, these represent a distribution whose variance is determined by the standard deviation. During training, points are randomly sampled from this distribution and passed to the decoder for reconstruction. The model is then trained not only to reconstruct inputs, but also to ensure that these distributions overlap and cover the latent space smoothly.20 Consequently, the latent space has no gaps since every point maps to a realistic output, and nearby points map to similar outputs.
At the same time, the decoder learns which combinations of features tend to occur together in the real data. If income and savings are positively correlated in the training data, for example, the decoder will produce profiles in which these features co-vary accordingly. In this way, VAEs implicitly capture the causal structure of input data.
It is this property that makes VAEs useful for generating counterfactual explanations. Continuing the example above, suppose that a bank uses an opaque model to assess loan-worthiness, and that Mary is denied a loan. Following the VAE approach, we would first use the encoder to compress her application into a point in the latent space. We could then use an optimization algorithm to search the latent space for a nearby point that, when decoded, produces a profile that the opaque model would accept rather than reject.21 Because the latent space is smooth, the nearby point the algorithm finds will produce a profile only slightly different from Mary’s. And because the decoder has learned which combinations of features are realistic, the resulting profile will also resemble that of a real applicant.
The key difference between the standard approach and the VAE approach is that in the latter, unlike the former, the search for counterfactuals takes place in the latent space rather than in the feature space. Recall that the standard approach searches for counterfactuals by perturbing the input features of Mary’s profile, treating each as causally independent. The VAE approach, by contrast, searches in latent space where the causal relationships between features have been implicitly learned by the model. When the optimization algorithm moves through the latent space, the decoder ensures that the resulting counterfactual respects those relationships. As such, this approach will not yield counterfactuals that suggest, for example, that Mary increase her savings by $10,000 without changing her salary. Accordingly, it seems that the VAE approach can avoid Karimi et al.’s objection.
Recent empirical work provides evidence that further supports this claim. Panagiotou et al. (2024), for example, introduce TABCF, a VAE-based method designed for the kind of tabular data used in many of the high-stakes decision-making contexts that Vredenburgh has in mind. They evaluate TABCF on five financial datasets and compare it against several alternatives, including the standard approach. They find that TABCF outperforms existing methods on several metrics, generating counterfactuals that more reliably flip the model’s decision and that require smaller changes to the original profile.22
Moreover, Panagiotou et al. identify a further problem with existing counterfactual methods. In particular, they find that the standard approach and its variants exhibit what they call feature-type bias: they disproportionately recommend changing numerical features, such as income or savings, while leaving categorical features, such as employment type or marital status, unchanged.23 TABCF avoids this bias because the VAE maps both numerical and categorical features into a smooth latent space. Accordingly, the VAE approach yields counterfactuals that better represent to decision-subjects the importance of changing categorical features.
This advantage of TABCF is directly relevant to the agential component of Vredenburgh’s right to explanation. If the standard approach recommends changing numerical features when changing categorical features may be less costly, for example, then decision-subjects will be provided with sub-optimal guidance. For reasons discussed above, this shortcoming suggests that we have a right to counterfactual explanations generated by the VAE approach rather than the standard approach.
It then seems that the VAE approach not only avoids the objection Karimi et al. raise against the standard approach but also improves upon other aspects of Wachter’s proposal. By searching for counterfactuals in a latent space that encodes the relationships between input features, the approach produces counterfactuals that correspond to applications that real people could actually have and, as Panagiotou et al. demonstrate, these explanations are superior to those generated using alternative counterfactual approaches on at least one metric. If this is true, counterfactuals generated using the VAE approach can plausibly satisfy the agency condition of the right to explanation better than alternatives. And if the VAE approach does, in fact, promote agency better than alternatives, then under the Scanlonian framework which Vredenburgh operates within, we would have a right to counterfactual explanations generated using this approach rather than others.
Even so, one might wonder whether the VAE approach requires knowledge that undermines its usefulness. In many applications, the dimensions of a VAE’s latent space do not automatically correspond to identifiable real-world properties. Ensuring that they do often requires careful architectural choices, such as constraining the model so that each latent dimension captures a single meaningful factor of variation. Without such choices, it may be unclear what changed between the original profile and the counterfactual, which would limit the explanation’s usefulness for agency.
However, for tabular data of the kind used in institutional decision-making, this concern is less pressing. In high-stakes contexts like loan decisions, hiring, and welfare eligibility, the relevant variables—income, debt, employment type, credit history—are well-documented and readily identifiable, making it straightforward for model architects to design the model in such a way that the latent variables will correspond to semantically-meaningful features. As such, the VAE approach can still plausibly protect our agential interests better than alternatives—even if we must first know some information about the features relevant to a decision. So, we would have a right to counterfactual explanations generated using the approach rather than those generated using the standard approach.24
Conclusion
I have argued that counterfactual explanations can fulfill the agential component of Vredenburgh’s right to explanation better than alternatives. Because VAEs learn to develop latent spaces that implicitly respect the causal relationships between input variables, they avoid the unrealistic counterfactuals that Karimi et al. believe plague the standard approach. The VAE approach thus provides decision-subjects with more actionable guidance about how to navigate institutional rules and achieve favorable outcomes.
This conclusion, however, requires qualification. Agency is only one component of informed self-advocacy. Representation and accountability impose further demands that counterfactual explanations—whether generated by VAEs or otherwise—likely cannot meet. As such, a complete fulfilment of the right to explanation will require explanations beyond what counterfactual methods can provide. So, if I am correct, I have just shown that the VAE approach constitutes an improvement over existing methods regarding one important dimension of Vredenburgh’s proposed right, and that the skepticism Grote and Paulo express about the feasibility of promoting agency through counterfactual explanation can, at least in part, be overcome.
Works Cited
Buckner, C. (2024). From Deep Learning to Rational Machines. Oxford University Press.
Grote, T., & Paulo, N. (2025). A minimalist account of the right to explanation. Philosophy & Technology, 38, 55. https://doi.org/10.1007/s13347-025-00888-3
Karimi, A. H., Schölkopf, B., & Valera, I. (2021). Algorithmic recourse: From counterfactual explanations to interventions. In Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (pp. 353–362). https://doi.org/10.1145/3442188.3445899
Kingma, D. P., & Welling, M. (2013). Auto-encoding variational Bayes. arXiv preprint arXiv:1312.6114.
Panagiotou, E., Heurich, M., Landgraf, T., & Ntoutsi, E. (2024). TABCF: Counterfactual explanations for tabular data using a transformer-based VAE. In Proceedings of the 5th ACM International Conference on AI in Finance (ICAIF ‘24), 1–9. https://doi.org/10.1145/3677052.3698642
Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence, 1(5), 206–215. https://doi.org/10.1038/s42256-019-0048-x
Vredenburgh, K. (2022). The right to explanation. The Journal of Political Philosophy, 30(2), 209–229. https://doi.org/10.1111/jopp.12262
Wachter, S., Mittelstadt, B., & Russell, C. (2018). Counterfactual explanations without opening the black box: Automated decisions and the GDPR. Harvard Journal of Law & Technology, 31(2), 841–887.
Footnotes
-
See e.g. Zerilli (2022) for pushback on this claim. ↩
-
In particular, Vredenburgh argues that institutional decision-makers ought to render these models “functionally transparent” in Creel (2020)‘s sense. In other words, they must give us a high-level description of how these models’ inputs relate to their outputs. Using simpler models to approximate the reasoning of opaque ones is one way of achieving such functional transparency. ↩
-
Vredenburgh 2022, p. 217. ↩
-
Ibid., p. 214. ↩
-
Ibid., p. 212. ↩
-
Ibid., p. 217. ↩
-
Grote and Paulo 2025, p. 7. ↩
-
Grote and Paulo additionally note that social environments are moving targets. Demographics, lifestyle choices, and consumer behavior can shift over time, and as a result, population-level causal generalizations drawn from the social science studies could be poor indicators of a model’s behavior. Such generalizations would mislead decision-subjects about what they ought to do to receive a desired outcome, thus inhibiting their agency. So, even if researchers can be expected to divide participants into fine-grained sub-groups, we still have reason to doubt that population-level causal generalizations can promote self-advocacy. ↩
-
Ibid., p. 6. ↩
-
It should be noted that Vredenburgh does not think agency is best enabled by model-to-model explanations; she thinks that such explanations better serve accountability and representation. But as Grote and Paulo note, there is no reason why these explanations could not, in principle, promote agency. ↩
-
Ibid., p. 10. ↩
-
Vredenburgh 2022, pp. 224-225. ↩
-
Grote and Paulo also argue that the free expert model has not proven effective in its original context. Vredenburgh takes anti-discrimination law as her blueprint, but as Grote and Paulo observe, even with legal assistance, applicants and employees face significant difficulties gathering evidence of unfair treatment because discriminatory acts tend to be subtle and covert. This objection is more directly relevant to the accountability component of informed self-advocacy than to agency, and so I set it aside here. ↩
-
Grote and Paulo do not directly discuss these problems; rather, they reference others who do. For our purposes, I will assess only one such reference. In particular, I will consider Karimi et al. (2021)‘s objection against current counterfactual methods as it is most relevant to agency. For other objections to counterfactual explanations, see Buijsman (2022) for a discussion of how existing counterfactual methods do not yield good explanations under a manipulationist account and Baron (2023) for an argument that existing counterfactual methods cannot achieve complete causal certification. ↩
-
One might worry, following Rudin (2019), that counterfactual explanations of opaque models are not genuine explanations at all. Rather than supplying some global description of a model’s internal logic, these explanations merely describe the model’s behavior in a particular region of their input space. Even if this concern is granted, it does not undermine my arguement or the usefulness of counterfactual explanations. What matters for the agency component of informed self-advocacy is just that a decision-subject knows what she can do to achieve a more favorable outcome. And a counterfactual that tells Mary, for example, what a realistic, minimally different accepted profile looks like can provide this information. Whether such guidance deserves the title of “explanation” in Rudin’s sense is a question I remain agnostic on here. ↩
-
Karimi et al., p.2. ↩
-
Ibid., p. 1. ↩
-
Vredenburgh remains intentionally ambiguous on what exactly agency requires. However, she suggests that the demandingness of the explanatory burden depends on the nature of the decision. A high-stakes decision that distributes harm—such as the decision to deny someone bail—would entail a high explanatory requirement. In cases like this, it is plausible that institutions have a requirement to provide optimal or near-optimal guidance on how decision-subjects can achieve a desired outcome, so the standard approach would be inadequate. However, only a small number of decisions carry such weight. As such, I will rely on other arguments to rebuff this objection. ↩
-
It should be noted that, although my upcoming argument bolsters the VAE approach, I am not committed to it. There may be other approaches to counterfactual explanation that outperform the VAE approach. I merely argue that, perhaps among other things, we have a right to the form of explanation that best promotes our agential interests, and currently, it seems that the VAE approach does so best. ↩
-
Buckner 2024, p. 222. ↩
-
Panagiotou et al. (2024) provide one such algorithm for generating counterfactual explanations using VAEs. Their algorithm is optimized to satisfy three desiderata: the decoded profile must flip the model’s decision, it must remain close in latent space to the original input, and it must change as few features as possible. ↩
-
Panagiotou et al., p.6. ↩
-
Ibid., p.4. ↩
-
Of course, future empirical research may unveil further objections to the VAE approach. For instance, it remains understudied how useful decision-subjects actually find counterfactual explanations generated using the VAE approach in real-life, institutional contexts. As such, I will reiterate that my aim is not primarily to bolster the VAE approach, but rather to demonstrate that recent developments in counterfactual explanation can, contrary to the skepticism expressed by Grote and Paulo, fulfill the agential component of the right to explanation. Whether VAEs will remain the dominant approach, and thus protect the interest better than alternatives, remains an open empirical question. ↩