<?xml version="1.0" encoding="UTF-8" standalone="no"?>
<!DOCTYPE article PUBLIC "-//NLM//DTD Journal Publishing DTD v2.3 20070202//EN" "journalpublishing.dtd">
<article xml:lang="EN" xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink" article-type="research-article">
<front>
<journal-meta>
<journal-id journal-id-type="publisher-id">Front. Artif. Intell.</journal-id>
<journal-title>Frontiers in Artificial Intelligence</journal-title>
<abbrev-journal-title abbrev-type="pubmed">Front. Artif. Intell.</abbrev-journal-title>
<issn pub-type="epub">2624-8212</issn>
<publisher>
<publisher-name>Frontiers Media S.A.</publisher-name>
</publisher>
</journal-meta>
<article-meta>
<article-id pub-id-type="doi">10.3389/frai.2024.1346700</article-id>
<article-categories>
<subj-group subj-group-type="heading">
<subject>Artificial Intelligence</subject>
<subj-group>
<subject>Original Research</subject>
</subj-group>
</subj-group>
</article-categories>
<title-group>
<article-title>Causal contextual bandits with one-shot data integration</article-title>
</title-group>
<contrib-group>
<contrib contrib-type="author" corresp="yes">
<name><surname>Subramanian</surname> <given-names>Chandrasekar</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<xref ref-type="corresp" rid="c001"><sup>&#x0002A;</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/2543115/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-original-draft/"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/data-curation/"/>
<role content-type="https://credit.niso.org/contributor-roles/investigation/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/software/"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
</contrib>
<contrib contrib-type="author">
<name><surname>Ravindran</surname> <given-names>Balaraman</given-names></name>
<xref ref-type="aff" rid="aff1"><sup>1</sup></xref>
<xref ref-type="aff" rid="aff2"><sup>2</sup></xref>
<uri xlink:href="http://loop.frontiersin.org/people/549562/overview"/>
<role content-type="https://credit.niso.org/contributor-roles/writing-review-editing/"/>
<role content-type="https://credit.niso.org/contributor-roles/funding-acquisition/"/>
<role content-type="https://credit.niso.org/contributor-roles/conceptualization/"/>
<role content-type="https://credit.niso.org/contributor-roles/methodology/"/>
<role content-type="https://credit.niso.org/contributor-roles/supervision/"/>
<role content-type="https://credit.niso.org/contributor-roles/validation/"/>
</contrib>
</contrib-group>
<aff id="aff1"><sup>1</sup><institution>Robert Bosch Center for Data Science and Artificial Intelligence, Indian Institute of Technology Madras</institution>, <addr-line>Chennai</addr-line>, <country>India</country></aff>
<aff id="aff2"><sup>2</sup><institution>Department of Computer Science and Engineering, Indian Institute of Technology Madras</institution>, <addr-line>Chennai</addr-line>, <country>India</country></aff>
<author-notes>
<fn fn-type="edited-by"><p>Edited by: Mohan Sridharan, University of Edinburgh, United Kingdom</p></fn>
<fn fn-type="edited-by"><p>Reviewed by: Alberto Ochoa Zezzatti, Universidad Aut&#x000F3;noma de Ciudad Ju&#x000E1;rez, Mexico</p>
<p>Reza Shahbazian, University of Calabria, Italy</p>
<p>Muhammad Yousuf Jat Baloch, Jilin University, China</p>
<p>Anam Nigar, Changchun University of Science and Technology, China</p></fn>
<corresp id="c001">&#x0002A;Correspondence: Chandrasekar Subramanian <email>sekarnet&#x00040;gmail.com</email></corresp>
</author-notes>
<pub-date pub-type="epub">
<day>06</day>
<month>12</month>
<year>2024</year>
</pub-date>
<pub-date pub-type="collection">
<year>2024</year>
</pub-date>
<volume>7</volume>
<elocation-id>1346700</elocation-id>
<history>
<date date-type="received">
<day>29</day>
<month>11</month>
<year>2023</year>
</date>
<date date-type="accepted">
<day>19</day>
<month>11</month>
<year>2024</year>
</date>
</history>
<permissions>
<copyright-statement>Copyright &#x000A9; 2024 Subramanian and Ravindran.</copyright-statement>
<copyright-year>2024</copyright-year>
<copyright-holder>Subramanian and Ravindran</copyright-holder>
<license xlink:href="http://creativecommons.org/licenses/by/4.0/"><p>This is an open-access article distributed under the terms of the Creative Commons Attribution License (CC BY). The use, distribution or reproduction in other forums is permitted, provided the original author(s) and the copyright owner(s) are credited and that the original publication in this journal is cited, in accordance with accepted academic practice. No use, distribution or reproduction is permitted which does not comply with these terms.</p></license>
</permissions>
<abstract>
<p>We study a contextual bandit setting where the agent has access to causal side information, in addition to the ability to perform multiple targeted experiments corresponding to potentially different context-action pairs&#x02014;simultaneously in one-shot within a budget. This new formalism provides a natural model for several real-world scenarios where parallel targeted experiments can be conducted and where some domain knowledge of causal relationships is available. We propose a new algorithm that utilizes a novel entropy-like measure that we introduce. We perform several experiments, both using purely synthetic data and using a real-world dataset. In addition, we study sensitivity of our algorithm&#x00027;s performance to various aspects of the problem setting. The results show that our algorithm performs better than baselines in all of the experiments. We also show that the algorithm is sound; that is, as budget increases, the learned policy eventually converges to an optimal policy. Further, we theoretically bound our algorithm&#x00027;s regret under additional assumptions. Finally, we provide ways to achieve two popular notions of fairness, namely counterfactual fairness and demographic parity, with our algorithm.</p></abstract>
<kwd-group>
<kwd>causality</kwd>
<kwd>fairness</kwd>
<kwd>causal contextual bandits</kwd>
<kwd>causal bandits</kwd>
<kwd>contextual bandit algorithm</kwd>
</kwd-group>
<counts>
<fig-count count="11"/>
<table-count count="1"/>
<equation-count count="39"/>
<ref-count count="34"/>
<page-count count="17"/>
<word-count count="11081"/>
</counts>
<custom-meta-wrap>
<custom-meta>
<meta-name>section-at-acceptance</meta-name>
<meta-value>Machine Learning and Artificial Intelligence</meta-value>
</custom-meta>
</custom-meta-wrap>
</article-meta>
</front>
<body>
<sec sec-type="intro" id="s1">
<title>1 Introduction</title>
<p>Learning to make decisions that depend on context has a wide range of applications&#x02014;software product experimentation, personalized medical treatments, recommendation systems, marketing campaign design, etc. Contextual bandits (Lattimore and Szepesv&#x000E1;ri, <xref ref-type="bibr" rid="B14">2020</xref>) have been used to model such problems with good success (Liu et al., <xref ref-type="bibr" rid="B15">2018</xref>; Sawant et al., <xref ref-type="bibr" rid="B22">2018</xref>; Bouneffouf et al., <xref ref-type="bibr" rid="B3">2020</xref>; Ameko et al., <xref ref-type="bibr" rid="B2">2020</xref>). In the contextual bandit framework, the agent interacts with an environment to learn a near-optimal policy that maps a context space to an action space.<xref ref-type="fn" rid="fn0001"><sup>1</sup></xref></p>
<p>A major challenge in wider application of contextual bandits to real world problems is the need for a large number samples, which is often prohibitively costly to obtain; indeed, this is a challenge with reinforcement learning in general (Dulac-Arnold et al., <xref ref-type="bibr" rid="B5">2021</xref>). Typically, this is mitigated by considering special cases and exploiting the structures present in those settings to obtain better algorithms. Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>) provided the first approach (and the only one so far) for contextual bandits where a causal graph was present as side information, in addition to the agent having the ability to perform interventions targeted on subgroups.</p>
<p>Specifically, in many real world scenarios, we often have some causal side information available from domain knowledge. For instance, in software product experimentation, we might know that <monospace>os</monospace> has a causal effect on <monospace>browser</monospace>, but not the other way around. Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>) showed that exploiting this causal knowledge can yield better contextual bandit policies.<xref ref-type="fn" rid="fn0002"><sup>2</sup></xref> They also introduced the notion of targeted interventions, which are interventions targeted on a specific subgroup specified by a particular assignment of values to the context variables.</p>
<p>This work proposes and studies a causal contextual bandits framework that has similarities with the one proposed in Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>) where the agent has access to causal side information and can perform targeted interventions; the agent&#x00027;s objective is to minimize <italic>simple regret</italic> (Lattimore and Szepesv&#x000E1;ri, <xref ref-type="bibr" rid="B14">2020</xref>). However, in contrast to the purely interactive setting in their work, we consider a setting where <italic>multiple targeted interventions can be performed simultaneously in one-shot</italic> at a cost. This, as we will see, fundamentally changes the agent&#x00027;s optimization problem due to the shared causal graph across the different interventions. However, importantly, it also opens up causal contextual bandits to new areas of application where it is possible to acquire additional data in one-shot at a cost and within a budget; some examples are discussed in Section 1.1. Section 1.2 provides a non-mathematical description of our framework, while Section 2.1 provides a mathematical specification of the problem. Section 1.4 provides a comparison of this work with related work.</p>
<p>To the best of our knowledge, there has not been any investigation on what the best way is for causal contextual bandit agents is to obtain additional experimental data in one shot and incorporate it with aim of learning a good policy.</p>
<sec>
<title>1.1 Motivating examples</title>
<p>In software product development, product teams are frequently interested in learning the best software variant to show or roll out to each subgroup of users. To achieve this, it is often possible to conduct targeting at scale <italic>simultaneously</italic>, i.e., in <italic>one shot</italic>, for various combinations of contexts (representing user groups) and actions (representing software variants)&#x02014;instead of one experiment at a time.<xref ref-type="fn" rid="fn0003"><sup>3</sup></xref> These targeted experiments can be used to compute relevant metrics (e.g., click-through rate) for each context-action pair. Further, we might have some qualitative domain knowledge of how some context variables are causally related to others; for example, we might know that <monospace>os</monospace> has a causal relationship to <monospace>browser</monospace>; we would like to exploit this knowledge to learn better policies. This can be naturally modeled as a question of &#x0201C;given the causal graph, what <italic>table</italic> of targeted experimental data needs to be acquired and how to integrate that data, so as to learn a good policy?&#x0201D; here the table&#x00027;s rows and columns are contexts and actions, and each cell specifies the number of data samples (zero or more) required corresponding to the context-action pair. At the end of the above training phase, the agent moves to an evaluation phase (e.g., when it is deployed), where it observes contexts, and takes actions using the learned policy.</p>
<p>As another example, in online marketing campaign strategy, the objective is to learn the best campaign to show to each group of people.<xref ref-type="fn" rid="fn0004"><sup>4</sup></xref> This is done by conducting experiments where different groups of people are shown different marketing campaigns simultaneously, and learning from the resultant outcomes. Further, letting context variables model the features of groups, we might have some knowledge about how some of the features are causally related to others; for example, we might know that <monospace>country</monospace> causally affects <monospace>income</monospace>. Our framework provides a natural model for this setting. Further, we might also be interested in ensuring that the campaigns meet certain criteria for fairness. For example, there might be some variables such as <monospace>race</monospace> that could be sensitive from a fairness perspective, and we would not want the agent to learn policies that depends on these variables. We discuss fairness implications of our algorithm in Section 5. These are just two of many scenarios where this framework provides a natural model. Two additional examples include experimental design for ads<xref ref-type="fn" rid="fn0005"><sup>5</sup></xref> and recommendation systems.<xref ref-type="fn" rid="fn0006"><sup>6</sup></xref></p>
<sec>
<title>1.1.1 Remark</title>
<p>The ideas and approach provided in this paper are not restricted to the examples mentioned above, and can be more generally applied by the wider scientific community. The scientific method, especially in the natural and social sciences, often involves performing experiments with real world entities such as people, animals or objects. Whenever there are opportunities for conducting multiple experiments in parallel (e.g., multiple people conducting social interventions at the same time on different groups) and there is some knowledge of causal relationships between the attributes of those entities (e.g., relationships between demographic attributes of beneficiaries), the framework and approach mentioned in this paper can be considered. The algorithm in this paper provides an efficient approach for designing the parallel experiments in these settings. However, the ethical aspects of such experiments have to be accounted for, before using in practice in such cases.</p>
</sec>
</sec>
<sec>
<title>1.2 Our framework</title>
<p>Our framework captures the various complexities and nuances described in the examples in Section 1.1. We present an overview here; please see Section 2.1 for the mathematical formalism. The agent&#x00027;s interactions consist of a learning phase where the agent incurs no regret, followed by an evaluation phase where regret is measured; this is also called <italic>simple regret</italic> (Lattimore and Szepesv&#x000E1;ri, <xref ref-type="bibr" rid="B14">2020</xref>). At the start of the learning phase, the agent is given a (possibly empty) log of offline data&#x02014;consisting of context-action-reward tuples&#x02014;generated from some unknown policy. The context variables are partitioned into two sets&#x02014;the main set and the auxiliary set (possibly empty). The agent observes all context variables during this phase, but learns a policy that only depends on the main set of context variables; this also provides a way to ensure that the learned policy meets certain definitions of fairness (see Section 5 for a more detailed discussion). Further, the agent also has some qualitative<xref ref-type="fn" rid="fn0007"><sup>7</sup></xref> causal side-information available, likely from domain knowledge. This causal side-information is encoded as a causal graph between contextual variables. A key implication of the causal graph is information leakage (Lattimore et al., <xref ref-type="bibr" rid="B13">2016</xref>; Subramanian and Ravindran, <xref ref-type="bibr" rid="B27">2022</xref>)&#x02014;getting samples for one context-action pair provides information about other context-action pairs because of shared pathways in the causal graph.</p>
<p>Given the logged data and the causal graph, the agent&#x00027;s problem is to decide the set of targeted experimental samples to acquire within a budget, and then integrate the returned samples. More specifically, the agent is allowed to make a <italic>one-shot</italic> request for data in the form of a table specifying the number of samples it requires for each context-action pair, subject to the total cost being within the budget. The environment then returns the requested samples after conducting the targeted interventions, and the agent integrates those samples to update its internal beliefs and learned policy. After this learning phase, the agent moves to an inference or evaluation phase, where it returns an action (according to the learned policy) for every context it encounters; its regret is measured at this point. The core problem of the agent is to choose these samples in a way that straddles the trade-off between choosing more samples for context-action pairs it knows is likely more valuable (given its beliefs) and choosing more samples to explore less-seen context-action pairs&#x02014;while taking into account the budget and the information leakage across all obtained samples arising from the causal graph.</p>
</sec>
<sec>
<title>1.3 Contributions</title>
<list list-type="order">
<list-item><p>This is the first work to study how to <italic>actively obtain and integrate</italic> a table of <italic>multiple samples in one-shot</italic> in a contextual bandit setting. Further, we study this in the presence of a causal graph, making it one of the very few works to study the <italic>utilization of causal side-information</italic> by contextual bandit agents. See Section 1.4 for a more detailed discussion on related work, and Section 2.1 for the mathematical formalism of the problem.</p></list-item>
<list-item><p>We propose a novel algorithm (Section 2.2.3) that works by minimizing a new entropy-like measure called &#x003A5;(.) that we introduce. See Section 2.2 for a full discussion on the approach.</p></list-item>
<list-item><p>We show results of extensive experiments using purely synthetically generated data and an experiment inspired by real-world data, that demonstrate that our algorithm performs better than baselines. We also study sensitivity of the results to key aspects of the problem setting. See Section 4.</p></list-item>
<list-item><p>We also show some theoretical results. Specifically, we show that the method is sound &#x02013; that is, as the budget tends to infinity, the algorithm&#x00027;s regret converges to 0 (Section 3.2). Further, we provide a bound on regret for a limited case (Section 3.1).</p></list-item>
<list-item><p>We discuss fairness implications of our method in Section 5. We show that it can achieve counterfactual fairness. Further, while the algorithm does not guarantee demographic parity, we provide a way to recover this notion of fairness, but with a reduction in performance.</p></list-item>
</list>
</sec>
<sec>
<title>1.4 Related work</title>
<p>Causal bandits have been studied in the last few years (e.g., Lattimore et al., <xref ref-type="bibr" rid="B13">2016</xref>; Yabe et al., <xref ref-type="bibr" rid="B31">2018</xref>; Lu et al., <xref ref-type="bibr" rid="B16">2020</xref>), but they study this in a multi-armed bandit setting where the problem is identification of one best action. There is only one work (Subramanian and Ravindran, <xref ref-type="bibr" rid="B27">2022</xref>) studying causal <italic>contextual</italic> bandits&#x02014;where the objective is to learn a <italic>policy</italic> mapping contexts to actions&#x02014;and this is the closest related work. While we do leverage some ideas introduced in that work in our methodology and in the design of experiments, our work differs fundamentally from this work in important ways. Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>) consider a standard interactive setting where the agent can repeatedly act, observe outcomes and update its beliefs, whereas in our work the agent has a one-shot data collection option for samples from <italic>multiple</italic> context-action pairs. This fundamentally changes the nature of the optimization problem as we will see in Section 2.2; it also makes it a more natural model in a different set of applications, some of which were discussed in Section 1.1. Further, our work allows for arbitrary costs for collecting those samples, whereas they assume every intervention is of equal cost.</p>
<p>Contextual bandits in purely offline settings, where decision policies are learned from logged data, is a well-studied problem. Most of the work involves inverse propensity weighting based methods (such as Swaminathan and Joachims, <xref ref-type="bibr" rid="B29">2015b</xref>,<xref ref-type="bibr" rid="B28">a</xref>; Joachims et al., <xref ref-type="bibr" rid="B10">2018</xref>). Contextual bandits are also well-studied in purely interactive settings (see Lattimore and Szepesv&#x000E1;ri, <xref ref-type="bibr" rid="B14">2020</xref> for a discussion on various algorithms). However, in contrast to our work, none of these methods can integrate causal side information or provide a way to actively acquire and integrate new targeted experimental data.</p>
<p>Active learning (Settles, <xref ref-type="bibr" rid="B24">2012</xref>) studies settings where an agent is allowed to query an oracle for ground truth labels for certain data points. This has been studied in supervised learning settings where the agent receives ground truth feedback; in contrast, in our case, the agent receives outcomes only for actions that were taken (&#x0201C;bandit feedback&#x0201D;). However, despite this difference, our approach can be viewed as incorporating some elements of active learning into contextual bandits by enabling the agent to acquire additional samples at a cost. There has been some work that has studied contextual bandits with costs and budget constraints (e.g., Agrawal and Goyal, <xref ref-type="bibr" rid="B1">2012</xref>; Wu et al., <xref ref-type="bibr" rid="B30">2015</xref>). There has also been work that has explored contextual bandit settings where the agent can not immediately integrate feedback from the environment, but can do so only in batches (Zhang et al., <xref ref-type="bibr" rid="B33">2022</xref>; Ren et al., <xref ref-type="bibr" rid="B20">2022</xref>; Han et al., <xref ref-type="bibr" rid="B9">2020</xref>). However, all these works consider settings where the samples are obtained through a standard contextual bandit interaction&#x02014;observe a context, choose an intervention, receive a reward; in contrast, our work considers <italic>targeted</italic> interventions where context values to determine the targeted subgroup is specified along with the intervention. Further, importantly, none of these works provide a way to integrate causal side information.</p>
</sec>
</sec>
<sec sec-type="methods" id="s2">
<title>2 Methodology</title>
<sec>
<title>2.1 Problem formalism</title>
<sec>
<title>2.1.1 Underlying model</title>
<p>We model the underlying environment as a causal model <inline-formula><mml:math id="M20"><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow></mml:math></inline-formula>, which is defined by a directed acyclic graph <inline-formula><mml:math id="M21"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula> over all variables (the &#x0201C;causal graph&#x0201D;) and a joint probability distribution &#x02119; that factorizes over <inline-formula><mml:math id="M22"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula> (Pearl, <xref ref-type="bibr" rid="B18">2009b</xref>; Koller and Friedman, <xref ref-type="bibr" rid="B11">2009</xref>). The set of variables in <inline-formula><mml:math id="M23"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula> consists of the action variable (<italic>X</italic>), the reward variable (<italic>Y</italic>), and the set of context variables (<inline-formula><mml:math id="M24"><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:math></inline-formula>). Each variable takes on a finite, known set of values; note that this is quite general, and accommodates categorical variables. <inline-formula><mml:math id="M25"><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:math></inline-formula> is partitioned into the set of main context variables <inline-formula><mml:math id="M26"><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula> and the set of (possibly empty) auxiliary context variables <inline-formula><mml:math id="M27"><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula>. That is, <inline-formula><mml:math id="M28"><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>&#x0222A;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>.</p>
<p>The agent knows only <inline-formula><mml:math id="M29"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula> but not <inline-formula><mml:math id="M30"><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow></mml:math></inline-formula>; therefore, the agent has no <italic>a priori</italic> knowledge of the conditional probability distributions (CPDs) of the variables.</p>
</sec>
<sec>
<title>2.1.2 Protocol</title>
<p>In addition to knowing <inline-formula><mml:math id="M31"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula>, the agent also has access to logged offline data, <inline-formula><mml:math id="M32"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>y</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>, where each (<bold>c</bold><sub><italic>i</italic></sub>, <italic>x</italic><sub><italic>i</italic></sub>, <italic>y</italic><sub><italic>i</italic></sub>) is sampled from <inline-formula><mml:math id="M33"><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow></mml:math></inline-formula> and <italic>x</italic><sub><italic>i</italic></sub> is chosen following some unknown policy. Unlike many prior works, such as Swaminathan and Joachims, <xref ref-type="bibr" rid="B29">2015b</xref>, the agent here does <italic>not</italic> have access to the logging propensities.</p>
<p>The agent then specifies in one shot the number of samples <inline-formula><mml:math id="M34"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula> it requires for each pair (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>).<xref ref-type="fn" rid="fn0008"><sup>8</sup></xref> We denote the full table of these values by <inline-formula><mml:math id="M35"><mml:mstyle mathvariant="bold"><mml:mtext>N</mml:mtext></mml:mstyle><mml:mo>&#x0225C;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x022C3;</mml:mo></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>. Given a (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>), there is an arbitrary cost <inline-formula><mml:math id="M36"><mml:mi>&#x003B2;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula> associated with obtaining those samples. The total cost should be at most a budget <italic>B</italic>. For each (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>), the environment returns <inline-formula><mml:math id="M37"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula> samples of the form <inline-formula><mml:math id="M38"><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>y</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x0007E;</mml:mo><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>Y</mml:mi><mml:mtext>&#x000A0;</mml:mtext><mml:mo>|</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mi>d</mml:mi><mml:mi>o</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula>.<xref ref-type="fn" rid="fn0009"><sup>9</sup></xref> Let&#x00027;s call this acquired dataset <inline-formula><mml:math id="M39"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>. The agent utilizes <inline-formula><mml:math id="M40"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> along with <inline-formula><mml:math id="M41"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to learn a good policy.</p>
</sec>
<sec>
<title>2.1.3 Objective</title>
<p>The agent&#x00027;s objective is to learn a policy <inline-formula><mml:math id="M42"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mtext>&#x000A0;</mml:mtext><mml:mo>:</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x02192;</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula> such that expected simple regret is minimized:</p>
<disp-formula id="E1a"><mml:math id="M43"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x0225C;</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>-</mml:mo><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where &#x003D5;<sup>&#x0002A;</sup> is an optimal policy, <inline-formula><mml:math id="M44"><mml:msubsup><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msubsup><mml:mo>&#x0225C;</mml:mo><mml:mo>&#x1D53C;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>|</mml:mo><mml:mi>d</mml:mi><mml:mi>o</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M45"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003BC;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo>&#x0225C;</mml:mo><mml:mo>&#x1D53C;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>|</mml:mo><mml:mi>d</mml:mi><mml:mi>o</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p><xref ref-type="table" rid="T1">Table 1</xref> provides a summary of the key notation used in this paper.</p>
<table-wrap position="float" id="T1">
<label>Table 1</label>
<caption><p>Summary of key notation.</p></caption>
<table frame="box" rules="all">
<thead>
<tr style="background-color:#919498;color:#ffffff">
<th valign="top" align="left"><bold>Notation</bold></th>
<th valign="top" align="left"><bold>Meaning</bold></th>
</tr>
</thead>
<tbody>
<tr>
<td valign="top" align="left"><italic>X</italic></td>
<td valign="top" align="left">Action variable</td>
</tr> <tr>
<td valign="top" align="left"><italic>Y</italic></td>
<td valign="top" align="left">Reward variable</td>
</tr> <tr>
<td valign="top" align="left"><inline-formula><mml:math id="M1"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td valign="top" align="left">Set of main context variables and set of auxiliary context variables, respectively; so the set of all context variables is <inline-formula><mml:math id="M2"><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>&#x0222A;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>.</td>
</tr> <tr>
<td valign="top" align="left">Capital letters</td>
<td valign="top" align="left">A random variable; e.g., <italic>C</italic><sub>1</sub> or <italic>X</italic></td>
</tr> <tr>
<td valign="top" align="left">Small letters</td>
<td valign="top" align="left">A random variable&#x00027;s value; e.g., <italic>c</italic><sub>1</sub> or <italic>x</italic></td>
</tr> <tr>
<td valign="top" align="left">Small bold font</td>
<td valign="top" align="left">An assignment of values to a set of random variables; for example, <bold>c</bold> denotes a specific choice of values taken by variables in <inline-formula><mml:math id="M3"><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:math></inline-formula></td>
</tr> <tr>
<td valign="top" align="left"><inline-formula><mml:math id="M4"><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula>, <inline-formula><mml:math id="M5"><mml:mover accent="true"><mml:mrow><mml:mo>&#x1D53C;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:math></inline-formula></td>
<td valign="top" align="left">Estimate of distribution &#x02119; and expectation &#x1D53C; based on current beliefs</td>
</tr> <tr>
<td valign="top" align="left"><inline-formula><mml:math id="M6"><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>V</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="script">V</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula></td>
<td valign="top" align="left">Set of values taken by the variable <italic>V</italic>, and set of variables <inline-formula><mml:math id="M7"><mml:mrow><mml:mi mathvariant="script">V</mml:mi></mml:mrow></mml:math></inline-formula>, respectively.</td>
</tr> <tr>
<td valign="top" align="left"><inline-formula><mml:math id="M8"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:math></inline-formula></td>
<td valign="top" align="left">The learned policy and an optimal policy, respectively</td>
</tr> <tr>
<td valign="top" align="left">&#x003A5;</td>
<td valign="top" align="left">Entropy-like measure used in our algorithm; defined in <xref ref-type="disp-formula" rid="E3">Equation 3</xref></td>
</tr> <tr>
<td valign="top" align="left"><bold>pa</bold><sub><italic>V</italic></sub></td>
<td valign="top" align="left">Value of variables in <italic>PA</italic><sub><italic>V</italic></sub>, the parents of <italic>V</italic></td>
</tr> <tr>
<td valign="top" align="left"><inline-formula><mml:math id="M9"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula></td>
<td valign="top" align="left">Number of samples requested corresponding to <italic>X</italic> &#x0003D; <italic>x</italic> and <inline-formula><mml:math id="M10"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr> <tr>
<td valign="top" align="left"><inline-formula><mml:math id="M11"><mml:mi>&#x003B2;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula></td>
<td valign="top" align="left">Cost of acquiring <inline-formula><mml:math id="M12"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula> samples corresponding to <italic>X</italic> &#x0003D; <italic>x</italic> and <inline-formula><mml:math id="M13"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula></td>
</tr> <tr>
<td valign="top" align="left"><italic>B</italic></td>
<td valign="top" align="left">budget</td>
</tr> <tr>
<td valign="top" align="left"><inline-formula><mml:math id="M14"><mml:mstyle mathvariant="bold"><mml:mtext>a</mml:mtext></mml:mstyle><mml:mrow><mml:mo>&#x02329;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="script">B</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x0232A;</mml:mo></mml:mrow></mml:math></inline-formula></td>
<td valign="top" align="left">If <bold>a</bold> is an assignment of values to <inline-formula><mml:math id="M15"><mml:mrow><mml:mi mathvariant="script">A</mml:mi></mml:mrow></mml:math></inline-formula>, then <inline-formula><mml:math id="M16"><mml:mstyle mathvariant="bold"><mml:mtext>a</mml:mtext></mml:mstyle><mml:mrow><mml:mo>&#x02329;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="script">B</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x0232A;</mml:mo></mml:mrow></mml:math></inline-formula> is assignment of those values to respective variables <inline-formula><mml:math id="M17"><mml:mrow><mml:mi mathvariant="script">B</mml:mi></mml:mrow></mml:math></inline-formula>; <inline-formula><mml:math id="M18"><mml:mstyle mathvariant="bold"><mml:mtext>a</mml:mtext></mml:mstyle><mml:mrow><mml:mo>&#x02329;</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="script">B</mml:mi></mml:mrow></mml:mrow><mml:mo>&#x0232A;</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x02205;</mml:mi></mml:math></inline-formula> if <inline-formula><mml:math id="M19"><mml:mrow><mml:mi mathvariant="script">A</mml:mi></mml:mrow><mml:mo>&#x02229;</mml:mo><mml:mrow><mml:mi mathvariant="script">B</mml:mi></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x02205;</mml:mi></mml:math></inline-formula>.</td>
</tr></tbody>
</table>
</table-wrap>
</sec>
<sec>
<title>2.1.4 Assumptions</title>
<p>We assume that <italic>X</italic> has exactly one outgoing edge, <italic>X</italic> &#x02192; <italic>Y</italic>, in <inline-formula><mml:math id="M46"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula>. This is suitable to express a wide range of problems such as personalized treatments or software experimentation where the problem is to learn the best action under a context, but the action or treatment does not affect context variables. We also make a commonly-made assumption (see Guo et al., <xref ref-type="bibr" rid="B8">2020</xref>) that there are no unobserved confounders. Similar to Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>), we make an additional assumption that simplifies the factorization in Section 2.2.2: <inline-formula><mml:math id="M47"><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>C</mml:mi><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">&#x000A0;confounds&#x000A0;</mml:mtext></mml:mstyle><mml:msup><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mstyle class="text"><mml:mtext class="textrm" mathvariant="normal">&#x000A0;and&#x000A0;</mml:mtext></mml:mstyle><mml:mi>Y</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow><mml:mtext>&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x021D2;</mml:mo><mml:mtext>&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>C</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>; a <italic>sufficient</italic> condition for this to be true is if <inline-formula><mml:math id="M48"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> is ancestral.<xref ref-type="fn" rid="fn0010"><sup>10</sup></xref> This last assumption is a simplifying assumption and can be relaxed in the future.</p>
</sec>
</sec>
<sec>
<title>2.2 Solution approach</title>
<sec>
<title>2.2.1 Overall idea</title>
<p>In our approach, the agent works by maintaining beliefs<xref ref-type="fn" rid="fn0011"><sup>11</sup></xref> regarding every conditional probability distribution (CPD) in <inline-formula><mml:math id="M50"><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow></mml:math></inline-formula>. It first uses <inline-formula><mml:math id="M51"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> to update its initial CPD beliefs; this, in itself, makes use of information leakage provided the causal graph. It next needs to choose <inline-formula><mml:math id="M52"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>, which is the core problem. The key tradeoff facing the agent is the following: it needs to choose between allocating more samples to context-action pairs that it believes are more valuable and to context-action pairs that it knows less about. Unlike Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>), it cannot interactively choose and learn, but instead has to choose the whole <inline-formula><mml:math id="M53"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> in one shot&#x02014;necessitating the need to account for multiple overlapping information leakage pathways resulting from the multitude of samples. In addition, these samples have a cost to acquire, given by an arbitrary cost function, along with a total budget.</p>
<sec>
<title>2.2.1.1 Toward solving this</title>
<p>To achieve this, we define a novel function &#x003A5;(<bold>N</bold>) that captures a measure of overall entropy weighted by value. The idea is that minimizing &#x003A5; results in a good policy; that is, the agent&#x00027;s problem now becomes that of minimizing &#x003A5; subject to budget constraints. In Section 2.2.2, we formally define &#x003A5;(<bold>N</bold>) and provide some intuition. Later, we provide experimental support (see Section 4), along with some theoretical grounding to this intuition (see Section 3).</p>
</sec>
</sec>
<sec>
<title>2.2.2 The optimization problem</title>
<p>Determine <inline-formula><mml:math id="M54"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula> for each (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>) such that</p>
<disp-formula id="E2a"><mml:math id="M55"><mml:mrow><mml:mi>&#x003A5;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>N</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>is minimized, subject to</p>
<disp-formula id="E3a"><mml:math id="M56"><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:mi>&#x003B2;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x02264;</mml:mo><mml:mi>B</mml:mi></mml:mrow></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M57"><mml:mstyle mathvariant="bold"><mml:mtext>N</mml:mtext></mml:mstyle><mml:mo>&#x0225C;</mml:mo><mml:msub><mml:mrow><mml:mo>&#x022C3;</mml:mo></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>. We will next define &#x003A5;(<bold>N</bold>).</p>
<sec>
<title>2.2.2.1 Defining the objective function &#x003A5;(<bold>N</bold>)</title>
<p>The conditional distribution &#x02119;(<italic>V</italic>|<bold>pa</bold><sub><italic>V</italic></sub>) for any variable <italic>V</italic> is modeled as a categorical distribution whose parameters are sampled from a Dirichlet distribution (the belief distribution). That is, &#x02119;(<italic>V</italic>|<bold>pa</bold><sub><italic>V</italic></sub>) &#x0003D; <sans-serif>Cat</sans-serif>(<italic>V</italic>; <italic>b</italic><sub>1</sub>, ..., <italic>b</italic><sub><italic>r</italic></sub>), where (<italic>b</italic><sub>1</sub>, ..., <italic>b</italic><sub><italic>r</italic></sub>)&#x0007E;<sans-serif>Dir</sans-serif>(&#x003B8;<sub><italic>V</italic>|<sub><bold>pa</bold><sub><italic>V</italic></sub></sub></sub>), and &#x003B8;<sub><italic>V</italic>|<sub><bold>pa</bold><sub><italic>V</italic></sub></sub></sub> is a vector denoting the parameters of the Dirichlet distribution.</p>
<p>Actions in a contextual bandit setting can be interpreted as <italic>do</italic>() interventions on a causal model (Zhang and Bareinboim, <xref ref-type="bibr" rid="B32">2017</xref>; Lattimore et al., <xref ref-type="bibr" rid="B13">2016</xref>). Therefore, the reward <italic>Y</italic> when an agent chooses action <italic>x</italic> against context <bold>c</bold><sup><italic>A</italic></sup> can be thought of as being sampled according to &#x02119;[<italic>Y</italic>|<italic>do</italic>(<italic>x</italic>), <bold>c</bold><sup><italic>A</italic></sup>]. Under the assumptions described in Section 2.1.4, we can factorize as follows:</p>
<disp-formula id="E1"><label>(1)</label><mml:math id="M58"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mo>&#x1D53C;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>|</mml:mo><mml:mi>d</mml:mi><mml:mi>o</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mtable style="text-align:axis;" equalrows="false" columnlines="none none none none none none none none none" equalcolumns="false" class="array"><mml:mtr><mml:mtd><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02208;</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:mrow></mml:munder></mml:mstyle><mml:mstyle><mml:mrow><mml:mo stretchy="false">[</mml:mo></mml:mrow></mml:mstyle><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle><mml:mrow><mml:mo>&#x02329;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>P</mml:mi><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x0232A;</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0220F;</mml:mo></mml:mrow><mml:mrow><mml:mi>c</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mo>|</mml:mo><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle><mml:mrow><mml:mo>&#x02329;</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>P</mml:mi><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x0232A;</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mstyle><mml:mrow><mml:mo stretchy="false">]</mml:mo></mml:mrow></mml:mstyle></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Crucially, note that our beliefs about each CPD in <xref ref-type="disp-formula" rid="E1">Equation 1</xref> are affected by samples corresponding to multiple (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>) due to the shared terms in the factorization. To capture this, we construct a CPD-level uncertainty measure which we call <italic>Q</italic>(.):</p>
<disp-formula id="E5a"><mml:math id="M59"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>Q</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>V</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold'><mml:mtext>N</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x0225C;</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mo class="qopname">ln</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E2"><label>(2)</label><mml:math id="M60"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mstyle mathvariant="sans-serif"><mml:mtext>En</mml:mtext></mml:mstyle><mml:msup><mml:mrow><mml:mstyle mathvariant="sans-serif"><mml:mtext>t</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>V</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mstyle><mml:mrow><mml:mo stretchy="true">|</mml:mo></mml:mrow></mml:mstyle></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mrow><mml:mo>&#x02329;</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x0232A;</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>&#x02329;</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo>&#x0232A;</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo>&#x02329;</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x0232A;</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo>&#x02329;</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>&#x0232A;</mml:mo></mml:mrow></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Here <sans-serif>Ent</sans-serif><sup><italic>new</italic></sup> is defined in the same way as in Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>), which we reproduce here. Let the length of the vector &#x003B8;<sub><italic>V</italic>|<sub><bold>pa</bold><sub><italic>V</italic></sub></sub></sub> be <italic>r</italic>. Let &#x003B8;<sub><italic>V</italic>|<sub><bold>pa</bold><sub><italic>V</italic></sub></sub></sub>[<italic>i</italic>] denote the <italic>i</italic>&#x00027;th entry of &#x003B8;<sub><italic>V</italic>|<sub><bold>pa</bold><sub><italic>V</italic></sub></sub></sub>. We define an object called <sans-serif>Ent</sans-serif> that captures a measure of our knowledge of the CPD:</p>
<disp-formula id="E7a"><mml:math id="M61"><mml:mrow><mml:mstyle mathvariant="sans-serif"><mml:mtext>Ent</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>V</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x0225C;</mml:mo><mml:mo>-</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo class="qopname">ln</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>j</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>j</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>We then define</p>
<disp-formula id="E8a"><mml:math id="M62"><mml:mrow><mml:mstyle mathvariant="sans-serif"><mml:mtext>En</mml:mtext></mml:mstyle><mml:msup><mml:mrow><mml:mstyle mathvariant="sans-serif"><mml:mtext>t</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>n</mml:mi><mml:mi>e</mml:mi><mml:mi>w</mml:mi></mml:mrow></mml:msup><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>V</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x0225C;</mml:mo><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow></mml:mfrac><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>i</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mstyle mathvariant="sans-serif"><mml:mtext>Ent</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mstyle mathvariant="sans-serif"><mml:mtext>Cat</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M63"><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msubsup><mml:mrow><mml:mi>b</mml:mi></mml:mrow><mml:mrow><mml:mi>r</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x0007E;</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>Dir</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>-</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>i</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>i</mml:mi><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p>Finally, we construct &#x003A5;(<bold>N</bold>) as:</p>
<disp-formula id="E3"><label>(3)</label><mml:math id="M64"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x003A5;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>N</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x0225C;</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>V</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup><mml:mo>&#x0222A;</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:mi>Y</mml:mi></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:mrow></mml:munder></mml:mstyle><mml:mi>Q</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>V</mml:mi><mml:mo>|</mml:mo><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle><mml:mrow><mml:mo>&#x02329;</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x0232A;</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold'><mml:mtext>N</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mover accent="true"><mml:mrow><mml:mo>&#x1D53C;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle><mml:mrow><mml:mo>&#x02329;</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x0232A;</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec>
<sec>
<title>2.2.2.2 Intuition behind <italic>Q</italic>(.) and &#x003A5;(.)</title>
<p>Intuitively, <sans-serif>Ent</sans-serif><sup><italic>new</italic></sup> provides a measure of entropy if <italic>one</italic> additional sample corresponding to (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>) is obtained and used to update beliefs. <italic>Q</italic>(.) builds on it and captures the fact that the beliefs regarding any CPD &#x02119;[<italic>V</italic>|<bold>pa</bold><sub><italic>V</italic></sub>] can be updated using information leakage<xref ref-type="fn" rid="fn0012"><sup>12</sup></xref> from samples corresponding to <italic>multiple</italic> (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>); it does this by selecting the relevant (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>) pairs making use of the causal graph <inline-formula><mml:math id="M66"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula> and aggregating them. In addition, <italic>Q</italic>(.) also captures the fact that entropy reduces non-linearly with the number of samples. Finally, &#x003A5;(<bold>N</bold>) provides an aggregate (weighted) resulting uncertainty from choosing <inline-formula><mml:math id="M67"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula> samples of each (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>). The weighting in &#x003A5;(.) provides a way for the agent to relatively prioritize context-action pairs that are higher-value according to its beliefs.</p>
<p>For further intuition, note that <italic>Q</italic>(&#x02119;[<italic>V</italic>|<bold>pa</bold><sub><italic>V</italic></sub>], <bold>N</bold>) captures the resultant uncertainty in the agent&#x00027;s knowledge of &#x02119;[<italic>V</italic>|<bold>pa</bold><sub><italic>V</italic></sub>] if samples as specified by <bold>N</bold> are obtained and integrated into its beliefs. &#x003A5;(<bold>N</bold>) not only aggregates the CPD-level measure <italic>Q</italic>(.) to sum over all CPDs, but also adds a weighting factor <inline-formula><mml:math id="M68"><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mover accent="true"><mml:mrow><mml:mo>&#x1D53C;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle><mml:mrow><mml:mo>&#x02329;</mml:mo><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>&#x0232A;</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>. This term ensures that the algorithm does not overallocate samples to improve knowledge of CPDs that are less &#x0201C;valuable&#x0201D; (i.e., either the context is very unlikely to be seen, or the rewards are too low).</p>
</sec>
</sec>
<sec>
<title>2.2.3 Algorithm</title>
<p>The full learning algorithm, which we call <sans-serif>CoBA</sans-serif>, is given as Algorithm 1 (<xref ref-type="fig" rid="F1">Figure 1A</xref>). After learning, the algorithm for inferencing on any test context (i.e., returning the action for the given context) is given as Algorithm 2 (<xref ref-type="fig" rid="F1">Figure 1B</xref>). The core problem (Step 2 in Algorithm 1) is a nonlinear optimization problem with nonlinear constraints and integer variables. It can be solved using any of the various existing solvers; refer Section 2.2.3.1 for details on the solver used in our experiments.</p>
<fig id="F1" position="float">
<label>Figure 1</label>
<caption><p>The learning <bold>(A)</bold> and inference <bold>(B)</bold> phases of our algorithm <sans-serif>CoBA</sans-serif>. After the learning phase is complete, the agent can be deployed to perform inference as per the inference algorithm.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1346700-g0001.tif"/>
</fig>
<sec>
<title>2.2.3.1 Solving the optimization problem in Step 2 of Algorithm 1</title>
<p>In our experiments in Section 4, we use the <monospace>scipy.optimize.differential_evolution</monospace> solver<xref ref-type="fn" rid="fn0013"><sup>13</sup></xref> from the <monospace>scipy</monospace> Python library to solve the problem in Section 4.2. This solver implements differential evolution (Storn and Price, <xref ref-type="bibr" rid="B25">1997</xref>), which is an evolutionary algorithm which makes very few assumptions about the problem. The <monospace>scipy.optimize.differential_evolution</monospace> method is quite versatile; for example, it allows the specification of the objective function as a Python callable, and also allows arbitrary nonlinear constraints of type <monospace>scipy.optimize.NonlinearConstraint</monospace>. However, a practitioner can use any suitable optimization algorithm or heuristic to solve this problem.</p>
</sec>
</sec>
</sec>
</sec>
<sec id="s3">
<title>3 Theoretical results</title>
<p>We first show a regret bound for our algorithm under additional assumptions; we leave a more general bound for future work. In addition, we also show a soundness result &#x02013; without any assumptions &#x02013; guaranteeing convergence in the limit. Importantly, in Section 4, we perform extensive experiments where we relax the assumptions made for the regret bound and demonstrate our algorithm&#x00027;s empirical performance.</p>
<sec>
<title>3.1 Regret bound (under additional assumptions)</title>
<p>In Theorem 3.1, we are interested in the case where <italic>B</italic> is finite. This is of interest in practical settings where the budget is usually small. We prove a regret bound under additional assumptions (A2) which we describe below. Define <inline-formula><mml:math id="M69"><mml:mi>m</mml:mi><mml:mo>&#x0225C;</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">min</mml:mo></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder><mml:mi>&#x003C0;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula>, where &#x003C0; is the (unknown) logging policy that generated <inline-formula><mml:math id="M70"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>; and <inline-formula><mml:math id="M71"><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">V</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x0225C;</mml:mo><mml:mo>|</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="script">V</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:math></inline-formula>.</p>
<p>Theorem 3.1 (Regret bound). Under the additional assumptions (A2) mentioned below, for any 0 &#x0003C; &#x003B4; &#x0003C; 1, with probability &#x02265; 1 &#x02212; &#x003B4;,</p>
<disp-formula id="E10"><mml:math id="M72"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msqrt><mml:mrow><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>B</mml:mi><mml:mo>-</mml:mo><mml:mi>&#x003F5;</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo class="qopname">ln</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:msqrt></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where <inline-formula><mml:math id="M73"><mml:mi>&#x003F5;</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msqrt><mml:mrow><mml:mi>B</mml:mi><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>/</mml:mo><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula>, ignoring terms that are constant in <italic>B</italic>, <italic>m</italic>, &#x003B4;, <inline-formula><mml:math id="M74"><mml:mo>|</mml:mo><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:math></inline-formula> and the number of possible context-action pairs.</p>
<p><italic>Proof</italic>. The proof closely follows the regret bound proof in Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>) and adapts it to our setting. The purpose of the proof is to establish an upper bound on performance, and not to provide a tight bound. Bounding regret without these additional assumptions (A2) is left for future work.</p>
<p>First, we define the assumptions (A2) under which the theorem holds:</p>
<list list-type="order">
<list-item><p>There is some non-empty past logged data (<inline-formula><mml:math id="M75"><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x0003E;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>), and it was generated by an (unknown) policy &#x003C0; where every action has a non-zero probability of being chosen (&#x003C0;(<italic>x</italic>|<bold>c</bold><sup><italic>A</italic></sup>) &#x0003E; 0, &#x02200;<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>). The latter is a commonly made assumption, for example, in inverse-propensity weighting based methods.</p></list-item>
<list-item><p><inline-formula><mml:math id="M76"><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x02265;</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mi>B</mml:mi></mml:math></inline-formula>, for some constant &#x003B1; &#x0003E; 0. This is generally achievable in real world settings since we usually have fairly large logged datasets (or it is quite cheap to acquire logged data; for example, think of search logs), and for the bound to hold we technically can have a very small &#x003B1; as long as it is &#x0003E;0. We also assume that <italic>B</italic> is finite, as discussed earlier.</p></list-item>
<list-item><p>The cost function &#x003B2; is constant; without loss of generality, we let this constant be equal to 1. This is a common case in real world applications, especially when we do not have estimates of cost; in those cases, we typically assign a fixed cost to all targeted experiments.</p></list-item>
</list>
<sec>
<title>3.1.1 Expression for overall bound</title>
<p>First, note that <xref ref-type="disp-formula" rid="E3">Equation 3</xref> in Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>) remains the same even for our case. This is because it depends only on the factorization of &#x1D53C;[<italic>Y</italic>|<italic>do</italic>(<italic>x</italic>), <bold>c</bold><sup><italic>A</italic></sup>] (see <xref ref-type="disp-formula" rid="E1">Equation 1</xref> in the main paper) and on the fact that in the evaluation phase the agent uses expected parameters of the CPDs (derived from its learned beliefs) to return an action for a given context.</p>
<p>Therefore, suppose, with probability &#x02265; 1 &#x02212; &#x003B4;<sub><italic>X</italic>,<sub><bold>pa</bold><sub><italic>Y</italic></sub></sub></sub>,</p>
<disp-formula id="E11"><mml:math id="M77"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mo>|</mml:mo><mml:mo>&#x02200;</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>|</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>|</mml:mo><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo><mml:mo>&#x02264;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>and with probability &#x02265; 1 &#x02212; &#x003B4;<sub><italic>C</italic>|<sub><bold>pa</bold><sub><italic>C</italic></sub></sub></sub>,</p>
<disp-formula id="E12"><mml:math id="M78"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mo>&#x02200;</mml:mo><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:mi>c</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo><mml:mo>&#x02264;</mml:mo><mml:msub><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where the expressions for &#x003B4;<sub><italic>X</italic>,<sub><bold>pa</bold><sub><italic>Y</italic></sub></sub></sub>, &#x003B4;<sub><italic>C</italic>|<sub><bold>pa</bold><sub><italic>C</italic></sub></sub></sub>, &#x003F5;<sub><italic>X</italic>,<sub><bold>pa</bold><sub><italic>Y</italic></sub></sub></sub> and &#x003F5;<sub><italic>C</italic>|<sub><bold>pa</bold><sub><italic>C</italic></sub></sub></sub> will be derived later in this section.</p>
<p>Then with probability <inline-formula><mml:math id="M79"><mml:mo>&#x02265;</mml:mo><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>-</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:munder><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder><mml:msub><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:math></inline-formula>, for any given <bold>c</bold><sup><italic>A</italic></sup>,</p>
<disp-formula id="E4"><label>(4)</label><mml:math id="M80"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtext class="textrm" mathvariant="normal">Regret</mml:mtext><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mo>&#x1D53C;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>|</mml:mo><mml:mi>d</mml:mi><mml:mi>o</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mo>*</mml:mo></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mo>&#x1D53C;</mml:mo><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>|</mml:mo><mml:mi>d</mml:mi><mml:mi>o</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>a</mml:mi></mml:mrow><mml:mrow><mml:mi>a</mml:mi><mml:mi>l</mml:mi><mml:mi>g</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x02264;</mml:mo><mml:mn>2</mml:mn><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:mn>3</mml:mn><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where we define</p>
<disp-formula id="E14"><mml:math id="M81"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x0225C;</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></disp-formula>
<p>and</p>
<disp-formula id="E15"><mml:math id="M82"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msubsup><mml:mo>&#x0225C;</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:math></disp-formula>
</sec>
<sec>
<title>3.1.2 Expressions for &#x003B4;<sub><italic>C</italic>|<sub><bold>pa</bold><sub><italic>C</italic></sub></sub></sub> and &#x003F5;<sub><italic>C</italic>|<sub><bold>pa</bold><sub><italic>C</italic></sub></sub></sub></title>
<p>Denote <inline-formula><mml:math id="M83"><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">V</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x0225C;</mml:mo><mml:mo>|</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mrow><mml:mi mathvariant="script">V</mml:mi></mml:mrow></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:math></inline-formula>. Let <italic>L</italic><sub><bold>pa</bold><sub><italic>C</italic></sub></sub> be the number of samples in <inline-formula><mml:math id="M84"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> where <italic>PA</italic><sub><italic>C</italic></sub> &#x0003D; <bold>pa</bold><sub><italic>C</italic></sub>. Now, our starting estimate of <inline-formula><mml:math id="M85"><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>C</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula> using <inline-formula><mml:math id="M86"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is computed as <inline-formula><mml:math id="M87"><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula>. Since <inline-formula><mml:math id="M88"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is built by observing <italic>C</italic><sup><italic>A</italic></sup> according to the natural distribution and choosing <italic>X</italic> according to some (unknown) policy, the proof of Lemma A.1 in Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>) can be followed if we replace <italic>T</italic>&#x02032; by &#x003B1;<italic>B</italic> since <inline-formula><mml:math id="M89"><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x02265;</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mi>B</mml:mi></mml:math></inline-formula>.</p>
<p>Therefore, suppose, with probability at least <inline-formula><mml:math id="M90"><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, it is true that</p>
<disp-formula id="E16"><mml:math id="M91"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mo>&#x02200;</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x02265;</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mi>B</mml:mi><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>If the above event is true, then it is also true that with probability at least 1 &#x02212; &#x003B4;<sub><italic>C</italic>|<sub><bold>pa</bold><sub><italic>C</italic></sub></sub></sub>, it is true that</p>
<disp-formula id="E17"><mml:math id="M92"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mo>&#x02200;</mml:mo><mml:mi>c</mml:mi><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>c</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo><mml:mo>&#x02264;</mml:mo><mml:msqrt><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mi>B</mml:mi><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where</p>
<disp-formula id="E18"><mml:math id="M93"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt><mml:mo>,</mml:mo><mml:mtext>&#x000A0;&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x0220F;</mml:mo></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:math></disp-formula>
<p>Therefore, we have that</p>
<disp-formula id="E5"><label>(5)</label><mml:math id="M94"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mi>B</mml:mi><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>|</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec>
<sec>
<title>3.1.3 Expressions for &#x003B4;<sub><italic>X</italic>,<sub><bold>pa</bold><sub><italic>Y</italic></sub></sub></sub> and &#x003F5;<sub><italic>X</italic>,<sub><bold>pa</bold><sub><italic>Y</italic></sub></sub></sub></title>
<p>Let <italic>L</italic><sub><italic>x</italic>,<sub><bold>pa</bold><sub><italic>Y</italic></sub></sub></sub> be the number of samples in <inline-formula><mml:math id="M95"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> where (<italic>X, PA</italic><sub><italic>Y</italic></sub>) &#x0003D; (<italic>x</italic>, <bold>pa</bold><sub><italic>Y</italic></sub>). As before, recollect that our estimate of <inline-formula><mml:math id="M96"><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula> is computed as <inline-formula><mml:math id="M97"><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003B8;</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msubsup><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>/</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mn>2</mml:mn></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula>. Further, the mean of <italic>L</italic><sub><italic>x</italic>,<sub><bold>pa</bold><sub><italic>Y</italic></sub></sub></sub> is <italic>at least</italic> <inline-formula><mml:math id="M98"><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>&#x000B7;</mml:mo><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mi>m</mml:mi><mml:mo>&#x0003E;</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mi>B</mml:mi><mml:mi>m</mml:mi><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula>, where <inline-formula><mml:math id="M99"><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">min</mml:mo></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder><mml:mi>&#x003C0;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula> and &#x003C0; is the unknown logging policy. From our set of assumptions (A2), we have that &#x003C0;(<italic>x</italic>|<bold>c</bold><sup><italic>A</italic></sup>) &#x0003E; 0, &#x02200;<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>; therefore, <italic>m</italic> &#x0003E; 0.</p>
<p>Given this, the proof of Lemma A.2 in Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>) can be followed.</p>
<p>Therefore, suppose, with probability at least <inline-formula><mml:math id="M100"><mml:mn>1</mml:mn><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:math></inline-formula>, it is true that</p>
<disp-formula id="E20"><mml:math id="M101"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mo>&#x02200;</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>L</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>&#x02265;</mml:mo><mml:mi>&#x003B1;</mml:mi><mml:mi>B</mml:mi><mml:mi>m</mml:mi><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where</p>
<disp-formula id="E21"><mml:math id="M102"><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msubsup><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:mrow></mml:math></disp-formula>
<p>If the above event is true, then it is also true that with probability at least 1 &#x02212; &#x003B4;<sub><italic>X</italic>,<sub><bold>pa</bold><sub><italic>Y</italic></sub></sub></sub>, it is true that</p>
<disp-formula id="E22"><mml:math id="M103"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mo>&#x02200;</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:mo>|</mml:mo><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>=</mml:mo><mml:mn>1</mml:mn><mml:mo>|</mml:mo><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mo>&#x02264;</mml:mo><mml:msqrt><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mi>B</mml:mi><mml:mi>m</mml:mi><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Therefore, we have that</p>
<disp-formula id="E6"><label>(6)</label><mml:math id="M104"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:msub><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mi>B</mml:mi><mml:mi>m</mml:mi><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
</sec>
<sec>
<title>3.1.4 Final bound</title>
<p>Now, we can plug the <xref ref-type="disp-formula" rid="E5">Equations 5</xref>, <xref ref-type="disp-formula" rid="E6">6</xref> back into <xref ref-type="disp-formula" rid="E4">Equation 4</xref>, and following the same union bound trick as in Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>) and some algebra, we get that for any 0 &#x0003C; &#x003B4; &#x0003C; 1, with probability &#x02265; 1 &#x02212; &#x003B4;,</p>
<disp-formula id="E24"><mml:math id="M105"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x02264;</mml:mo><mml:mn>3</mml:mn><mml:msub><mml:mrow><mml:mo>&#x1D53C;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msqrt><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mi>m</mml:mi><mml:mi>B</mml:mi><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<disp-formula id="E7"><label>(7)</label><mml:math id="M106"><mml:mtable class="eqnarray" columnalign="left"><mml:mtr><mml:mtd><mml:mtext>&#x02003;</mml:mtext><mml:mo>&#x0002B;</mml:mo><mml:mn>3</mml:mn><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>C</mml:mi><mml:mo>&#x02208;</mml:mo><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:msub><mml:mrow><mml:mo>&#x1D53C;</mml:mo></mml:mrow><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msqrt><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn></mml:mrow><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mi>B</mml:mi><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>-</mml:mo><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mn>2</mml:mn><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where</p>
<disp-formula id="E26"><mml:math id="M107"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>C</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt><mml:mo>,</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:msubsup><mml:mrow><mml:mi>&#x003F5;</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi><mml:mo>,</mml:mo><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msubsup><mml:mo>=</mml:mo><mml:msqrt><mml:mrow><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:mi>&#x003B1;</mml:mi><mml:mi>B</mml:mi></mml:mrow><mml:mrow><mml:mn>2</mml:mn></mml:mrow></mml:mfrac></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mo>|</mml:mo><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow><mml:mo>|</mml:mo></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>It can be simplified as presented in <bold>Theorem 3.1</bold> as:</p>
<disp-formula id="E27"><mml:math id="M108"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>R</mml:mi><mml:mi>e</mml:mi><mml:mi>g</mml:mi><mml:mi>r</mml:mi><mml:mi>e</mml:mi><mml:mi>t</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msqrt><mml:mrow><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>m</mml:mi><mml:mi>B</mml:mi><mml:mo>-</mml:mo><mml:mi>&#x003F5;</mml:mi></mml:mrow></mml:mfrac></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo class="qopname">ln</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:msqrt></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>where</p>
<disp-formula id="E28"><mml:math id="M109"><mml:mrow><mml:mi>&#x003F5;</mml:mi><mml:mo>&#x02208;</mml:mo><mml:mi>O</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msqrt><mml:mrow><mml:mi>B</mml:mi><mml:mo class="qopname">ln</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>P</mml:mi><mml:msub><mml:mrow><mml:mi>A</mml:mi></mml:mrow><mml:mrow><mml:mi>Y</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x003B4;</mml:mi></mml:mrow></mml:mfrac></mml:mrow></mml:msqrt></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>This completes the proof.</p>
</sec>
<sec>
<title>3.1.5 Discussion</title>
<p>The bound is inversely related to <inline-formula><mml:math id="M110"><mml:msqrt><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msqrt></mml:math></inline-formula>, which is analogous to the inverse relation to <inline-formula><mml:math id="M111"><mml:msqrt><mml:mrow><mml:mi>T</mml:mi></mml:mrow></mml:msqrt></mml:math></inline-formula> that other algorithms (albeit for different, but related, settings) such as Lattimore et al. (<xref ref-type="bibr" rid="B13">2016</xref>) and Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>) have. The regret is inversely related to <italic>m</italic>, which can be interpreted as a measure of &#x0201C;hardness&#x0201D; of a problem (a larger <italic>m</italic> intuitively means the problem instance is easier); similar approaches to defining bounds in terms of hardness of the problem instance has been used in other such as Lattimore et al. (<xref ref-type="bibr" rid="B13">2016</xref>), Sen et al. (<xref ref-type="bibr" rid="B23">2017</xref>), and Yabe et al. (<xref ref-type="bibr" rid="B31">2018</xref>). The size of the space of possible interventions is <inline-formula><mml:math id="M112"><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:math></inline-formula>. The bound grows with <inline-formula><mml:math id="M113"><mml:msqrt><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:math></inline-formula>. However, note that <inline-formula><mml:math id="M114"><mml:mi>m</mml:mi><mml:mo>=</mml:mo><mml:munder class="msub"><mml:mrow><mml:mo class="qopname">min</mml:mo></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder><mml:mi>&#x003C0;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula>. This means, that even if the logging policy that generated <inline-formula><mml:math id="M115"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> was random (the best case), then <italic>m</italic> &#x0003D; 1/<italic>M</italic><sub><italic>X</italic></sub> and it forces another <inline-formula><mml:math id="M116"><mml:msqrt><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:msub></mml:mrow></mml:msqrt></mml:math></inline-formula> term into the bound, thereby making the bound grow at the rate of <inline-formula><mml:math id="M117"><mml:msqrt><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub><mml:mo class="qopname">ln</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mi>X</mml:mi></mml:mrow></mml:msub><mml:msub><mml:mrow><mml:mi>M</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msqrt></mml:math></inline-formula>; if <italic>m</italic> is smaller than 1/<italic>M</italic><sub><italic>X</italic></sub>, then the bound would grow faster.</p>
</sec>
</sec>
<sec>
<title>3.2 Soundness</title>
<p>Section 3.1 looked at the case where <italic>B</italic> is finite. In contrast, in this section, we consider the case where <italic>B</italic> &#x02192; &#x0221E;. <bold>Theorem 3.2</bold> demonstrates the soundness of our approach by showing that as the budget increases, the learned policy will eventually converge to an optimal policy.</p>
<p>Theorem 3.2 (Soundness). As <italic>B</italic> &#x02192; &#x0221E;, <italic>Regret</italic> &#x02192; 0.</p>
<p><italic>Proof</italic>. We would like to show that <italic>Regret</italic> &#x02192; 0 as <italic>B</italic> &#x02192; &#x0221E;. As <italic>B</italic> &#x02192; &#x0221E;, in the limit, the problem becomes unconstrained minimization of &#x003A5;(<bold>N</bold>). Note that for all <bold>N</bold>, &#x003A5;(<bold>N</bold>) &#x02265; 0. Therefore, the smallest possible value of &#x003A5;(<bold>N</bold>) is 0.</p>
<p>First, note that <inline-formula><mml:math id="M118"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo>&#x02192;</mml:mo><mml:mi>&#x0221E;</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x02200;</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mtext>&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x021D2;</mml:mo><mml:mtext>&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mi>&#x003A5;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>N</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x02192;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>. This is because &#x02200;(<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>),</p>
<disp-formula id="E29"><mml:math id="M119"><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo>&#x02192;</mml:mo><mml:mi>&#x0221E;</mml:mi><mml:mtext>&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x021D2;</mml:mo><mml:mtext>&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mo class="qopname">ln</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac><mml:mo>&#x02192;</mml:mo><mml:mn>0</mml:mn></mml:mrow></mml:math></disp-formula>
<p>which, in turn, makes <italic>Q</italic>(&#x02119;[<italic>V</italic>|<bold>pa</bold><sub><italic>V</italic></sub>], <bold>N</bold>) &#x02192; 0, &#x02200;(<italic>V</italic>, <bold>pa</bold><sub><italic>V</italic></sub>). From <xref ref-type="disp-formula" rid="E3">Equation 3</xref>, it is easy to see that this causes &#x003A5;(<bold>N</bold>) &#x02192; 0.</p>
<p>Also note that <inline-formula><mml:math id="M120"><mml:mi>&#x003A5;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>N</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x02192;</mml:mo><mml:mn>0</mml:mn><mml:mtext>&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:mo>&#x021D2;</mml:mo><mml:mtext>&#x000A0;&#x000A0;&#x000A0;</mml:mtext><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo>&#x02192;</mml:mo><mml:mi>&#x0221E;</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x02200;</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula>. To see this, let &#x003A5;(<bold>N</bold>) &#x02192; 0, and consider the case where there exists a (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>) such that <inline-formula><mml:math id="M121"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula> is finite. That means that there is at least one term of the form</p>
<disp-formula id="E30"><mml:math id="M122"><mml:mrow><mml:mfrac><mml:mrow><mml:mn>1</mml:mn></mml:mrow><mml:mrow><mml:mn>1</mml:mn><mml:mo>&#x0002B;</mml:mo><mml:mo class="qopname">ln</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo>&#x0002B;</mml:mo><mml:mn>1</mml:mn></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:mrow></mml:math></disp-formula>
<p>which occurs in the <italic>Q</italic>(.) function of at least one CPD &#x02119;[<italic>V</italic>|<bold>pa</bold><sub><italic>V</italic></sub>], causing <italic>Q</italic>(&#x02119;[<italic>V</italic>|<bold>pa</bold><sub><italic>V</italic></sub>], <bold>N</bold>) &#x0003E; 0 since <sans-serif>Ent</sans-serif><sup><italic>new</italic></sup> &#x0003E; 0. This, in turn, causes &#x003A5;(<bold>N</bold>) &#x02192; &#x00338;0, resulting in a contradiction. Note that this makes an implicit technical assumption that &#x02119;[<bold>c</bold>] &#x0003E; 0, &#x02200;<bold>c</bold> and that min(<sans-serif>val</sans-serif>(<italic>Y</italic>)) &#x0003E; 0; these are stronger assumptions than necessary, and could be weakened in the future.</p>
<p>Thus, <inline-formula><mml:math id="M123"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo>&#x02192;</mml:mo><mml:mi>&#x0221E;</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x02200;</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x021D4;</mml:mo><mml:mi>&#x003A5;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>N</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x02192;</mml:mo><mml:mn>0</mml:mn></mml:math></inline-formula>. In other words, each (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>) gets a number of samples tending toward infinity if and only if &#x003A5; tends to 0. Thus, since &#x003A5;(<bold>N</bold>) &#x02265; 0, Algorithm 1 will allocate <inline-formula><mml:math id="M124"><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mo>&#x02192;</mml:mo><mml:mi>&#x0221E;</mml:mi><mml:mo>,</mml:mo><mml:mo>&#x02200;</mml:mo><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula>. Each CPD has at least one (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>) whose samples will be used to update its beliefs in Algorithm 1 (due to the technical assumption mentioned above). This means that each CPD will have its beliefs updated a number of times approaching infinity. Thus, for any (<italic>V</italic>, <bold>pa</bold><sub><italic>V</italic></sub>), <inline-formula><mml:math id="M125"><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>V</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x02192;</mml:mo><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>V</mml:mi><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>pa</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>V</mml:mi></mml:mrow></mml:msub></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>. As a result, we have that <inline-formula><mml:math id="M126"><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>&#x02192;</mml:mo><mml:mi>&#x02119;</mml:mi></mml:math></inline-formula>.</p>
<p>Since the agent&#x00027;s policy constructed using &#x02119; will necessarily be optimal, we have that <italic>Regret</italic> &#x02192; 0. This completes the proof.</p>
</sec>
</sec>
<sec id="s4">
<title>4 Experimental results</title>
<sec>
<title>4.1 Baselines and experimental setup</title>
<sec>
<title>4.1.1 Baselines</title>
<p>There are no existing algorithms that directly map to our setting. Therefore, we construct a set of natural baselines and study the performance of our algorithm <sans-serif>CoBA</sans-serif> against them. These baselines cover standard strategies for allocation without a way to capture information leakage explicitly (which our algorithm exploits). <sans-serif>EqualAlloc</sans-serif> allocates an equal number of samples to all (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>); this provides a good distribution of samples to all context-action pairs. <sans-serif>MaxSum</sans-serif> maximizes the <italic>total</italic> number of samples summed over all (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>). <sans-serif>PropToValue</sans-serif> allocates a number of samples to (<italic>x</italic>, <bold>c</bold><sup><italic>A</italic></sup>) that is proportional to <inline-formula><mml:math id="M127"><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mover accent="true"><mml:mrow><mml:mo>&#x1D53C;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>|</mml:mo><mml:mi>d</mml:mi><mml:mi>o</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>; this allocates relatively more samples to context-action pairs that are more &#x0201C;valuable&#x0201D; based on the agent&#x00027;s current beliefs. All baselines first involve updating the agent&#x00027;s starting beliefs regarding the CPDs of <inline-formula><mml:math id="M128"><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow></mml:math></inline-formula> using <inline-formula><mml:math id="M129"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> (same as Step 1 of Algorithm 1) before allocating samples for active obtainment as detailed above. After <inline-formula><mml:math id="M130"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is returned by the environment, all baselines use it update their beliefs (same as Step 4 of Algorithm 1).</p>
</sec>
<sec>
<title>4.1.2 Experiments</title>
<p>Similar to Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>), we consider a causal model <inline-formula><mml:math id="M131"><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow></mml:math></inline-formula> whose causal graph <inline-formula><mml:math id="M132"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula> consists of the following edges: <italic>C</italic><sub>1</sub> &#x02192; <italic>C</italic><sub>0</sub>, <italic>C</italic><sub>0</sub> &#x02192; <italic>X, C</italic><sub>0</sub> &#x02192; <italic>Y, X</italic> &#x02192; <italic>Y</italic>. We let <inline-formula><mml:math id="M133"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mn>1</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula> and <inline-formula><mml:math id="M134"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mrow><mml:mo>{</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mi>C</mml:mi></mml:mrow><mml:mrow><mml:mn>0</mml:mn></mml:mrow></mml:msub></mml:mrow><mml:mo>}</mml:mo></mml:mrow></mml:math></inline-formula>.</p>
<p>The causal graph is illustrated in <xref ref-type="fig" rid="F2">Figure 2</xref>. We use this causal graph for all experiments except Experiment 3.</p>
<fig id="F2" position="float">
<label>Figure 2</label>
<caption><p>Causal graph used in all experiments except Experiment 3.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1346700-g0002.tif"/>
</fig>
<p>Experiments 1 and 2 analyze the performance of our algorithm in a variety of settings, similar to those used in Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>). Experiment 3 analyzes the performance of the algorithm on a setting calibrated using real-world CRM sales-data provided in Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>). For details of all the parameterizations, please refer to the <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>. Experiments 4 through 7 analyze sensitivity of our algorithm&#x00027;s performance to various aspects of the problem setting. In all experiments, except Experiment 6, we set the cost function &#x003B2;(.) to be proportional to the number of samples &#x02013; a natural definition of cost; in Experiment 6, we analyze sensitivity to cost function choice. Section 4.6 reports results of Experiment 1 and 2 for larger values of <italic>B</italic> (until all algorithms converge), providing empirical evidence of our algorithm&#x00027;s improved <italic>asymptotic behavior</italic>. Further experiments providing more insights into <italic>why</italic> our algorithm performs better than baselines are discussed in Section 4.7. A few additional experiments for intuition are provided in the <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>.</p>
</sec>
<sec>
<title>4.1.3 Remark</title>
<p>If the specific parameterization of <inline-formula><mml:math id="M135"><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow></mml:math></inline-formula> were given <italic>a priori</italic>, it is possible to come up with an algorithm that performs optimally in that particular setting. However, the objective is to design a method that performs well <italic>overall</italic> without this <italic>a priori</italic> information. Consider the relative performance of the baselines in Experiments 2 and 3. We will see that while <sans-serif>EqualAlloc</sans-serif> performs better than <sans-serif>MaxSum</sans-serif> and <sans-serif>PropToValue</sans-serif> in Experiments 3 (Section 4.4), it performs worse than those two in Experiment 2 (Section 4.3). However, our algorithm performs better than all three baselines in all experiments, corroborating our algorithm&#x00027;s overall better performance.</p>
</sec>
</sec>
<sec>
<title>4.2 Experiment 1 (representative settings)</title>
<p>Different parameterizations of <inline-formula><mml:math id="M136"><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow></mml:math></inline-formula> can produce a wide range of possible settings. Given this, the first experiment studies the performance of our algorithm over a set of &#x0201C;<italic>representative settings</italic>.&#x0201D; Each of these settings has a natural interpretation; for example, <inline-formula><mml:math id="M137"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> could represent the set of person-level features that we are learning a personalized treatment for, or it could represent the set of customer attributes over which we&#x00027;re learning a marketing policy. The settings capture the intuition that high-value contexts (contexts for which, if the optimal action is learned, high expected rewards accrue to the agent) occur relatively less frequently (say, 20% of the time), but that there can be variation in other aspects. Specifically, the variations come from the number of different values of <bold>c</bold><sup><italic>A</italic></sup> over which the 20% probability mass is spread, and in how &#x0201C;risky&#x0201D; a particular context is (e.g., difference in rewards between the best and worst actions). For details of the parameterizations, please refer to the <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>. The number of samples in the initial dataset <inline-formula><mml:math id="M138"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is kept at <inline-formula><mml:math id="M139"><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>5</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:mo>|</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo><mml:mo>&#x000B7;</mml:mo><mml:mo>|</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:math></inline-formula>. We consider a uniformly exploring logging policy for <inline-formula><mml:math id="M140"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula>; that is, context variables for each sample are realized as per the natural distribution induced by <inline-formula><mml:math id="M141"><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow></mml:math></inline-formula>, but <italic>X</italic> is chosen randomly. In each run, the agent is presented with a randomly selected setting from the representative set. Results are averaged over 50 independent runs; error bars display &#x000B1;2 standard errors.</p>
<p><xref ref-type="fig" rid="F3">Figure 3</xref> provides the results of this Experiment. It plots the value of regret (normalized to [0, 1] since different settings have different ranges for regret) as budget <italic>B</italic> increases. We see that our algorithm performs better than all baselines. Our algorithm also retains its relatively lower regret at all values of <italic>B</italic>, providing empirical evidence of overall better regret performance.</p>
<fig id="F3" position="float">
<label>Figure 3</label>
<caption><p>Mean regrets normalized to [0, 1] for Experiment 1 (Section 4.2). Our algorithm performs better than all baselines, with gap in performance being higher for smaller budgets.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1346700-g0003.tif"/>
</fig>
</sec>
<sec>
<title>4.3 Experiment 2 (randomized parameters)</title>
<p>To ensure that the results are not biased due to our choice of the representative set in Experiment 1, this experiment studies the performance of our algorithm when we <italic>directly randomize the parameters</italic> of the CPDs in each run, subject to realistic constraints. Specifically, in each run, we (1) randomly pick an <italic>i</italic> &#x02208; {1, ..., &#x0230A;|<sans-serif>val</sans-serif>(<italic>C</italic><sub>1</sub>)|/2&#x0230B;}, (2) distribute 20% of the probability mass randomly over the smallest <italic>i</italic> values of <italic>C</italic><sub>1</sub>, and (3) distribute the remaining 80% of the mass over the remaining values of <italic>C</italic><sub>1</sub>. The smallest <italic>i</italic> values of <italic>C</italic><sub>1</sub> have higher value (i.e., the agent obtains higher rewards when the optimal action is chosen) than the other <italic>C</italic><sub>1</sub> values. Intuitively, this captures the commonly observed 80&#x02013;20 pattern (for example, 20% of the customers often contribute to around 80% of the revenue); but we randomize the other aspects. For details of all the parameterizations, please refer to the <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>. Averaging over runs provides an estimate of the performance of the algorithms on expectation. The number of samples in the initial dataset <inline-formula><mml:math id="M142"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is kept at <inline-formula><mml:math id="M143"><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>25</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:mo>|</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo><mml:mo>&#x000B7;</mml:mo><mml:mo>|</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:math></inline-formula>. The results are averaged over 50 independent runs; error bars display &#x000B1;2 standard errors.</p>
<p><xref ref-type="fig" rid="F4">Figure 4</xref> shows that our algorithm performs better than all baselines in this experiment. Our algorithm also demonstrates overall better regret performance by achieving the lowest regret for every choice of <italic>B</italic>.</p>
<fig id="F4" position="float">
<label>Figure 4</label>
<caption><p>Mean regrets for Experiment 2 (Section 4.3). Our algorithm performs better than all baselines.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1346700-g0004.tif"/>
</fig>
</sec>
<sec>
<title>4.4 Experiment 3 (calibrated using real-world data)</title>
<p>While Experiments 1 and 2 study purely synthetic settings, this experiment seeks to study the performance of our algorithm in <italic>realistic scenarios</italic>. We use the same causal graph used in the real world-inspired experiment in Section 4.2 of Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>) and calibrate the CPDs using the data provided there. For parameterizations, refer to the <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>.</p>
<p>The objective is to learn a policy that can assist salespeople by learning to decide how many outgoing calls to make in an ongoing deal, given just the type of deal and size of customer, so as to maximize a reward metric. The variables are related to each other causally as per the causal graph [Figure 3a in Subramanian and Ravindran (<xref ref-type="bibr" rid="B27">2022</xref>)]. The number of samples in the initial dataset <inline-formula><mml:math id="M144"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> is kept at <inline-formula><mml:math id="M145"><mml:mn>0</mml:mn><mml:mo>.</mml:mo><mml:mn>125</mml:mn><mml:mo>&#x000B7;</mml:mo><mml:mo>|</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo><mml:mo>&#x000B7;</mml:mo><mml:mo>|</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:math></inline-formula>. The results are averaged over 50 independent runs; error bars display &#x000B1;2 standard errors. <xref ref-type="fig" rid="F5">Figure 5</xref> shows the results of the experiment. Our algorithm performs better than all other algorithms in this real-world inspired setting as well. Further, it retains its better performance at every value of <italic>B</italic>.</p>
<fig id="F5" position="float">
<label>Figure 5</label>
<caption><p>Results of the real-world data inspired experiment (Experiment 3, Section 4.4). Our algorithm achieves better mean regrets than baselines.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1346700-g0005.tif"/>
</fig>
</sec>
<sec>
<title>4.5 Experiments 4 through 7 (robustness to problem setting)</title>
<p>Experiments 4 through 7 study <italic>sensitivity of the results to key aspects</italic> that define our settings. To aid this analysis, instead of regret, we consider a more aggregate measure which we call AUC. For any run, AUC is computed for a given algorithm by summing over <italic>B</italic> the regrets for that algorithm; this provides an approximation of the area under the curve (hence the name). We then study the sensitivity of AUC to various aspects of the setting or environment.</p>
<sec>
<title>4.5.1 Experiment 4 (&#x0201C;narrowness&#x0201D; of <inline-formula><mml:math id="M146"><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow></mml:math></inline-formula>)</title>
<p>We use the term &#x0201C;narrowness&#x0201D; informally. Since our algorithm <sans-serif>CoBA</sans-serif> exploits the information leakage in the causal graph, we expect it to achieve better performance when there is more leakage. To see this, suppose we do a forward sampling (Koller and Friedman, <xref ref-type="bibr" rid="B11">2009</xref>) of <inline-formula><mml:math id="M147"><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow></mml:math></inline-formula>; then, intuitively, more leakage occurs when more samples require sampling overlapping CPDs. For this experiment, we proxy this by varying |<sans-serif>val</sans-serif>(<italic>C</italic><sub>0</sub>)| while keeping |<sans-serif>val</sans-serif>(<italic>C</italic><sub>1</sub>)| fixed. The rest of the setting is the same as in Experiment 2. A lower |<sans-serif>val</sans-serif>(<italic>C</italic><sub>0</sub>)| means that the causal model is more &#x0201C;squeezed&#x0201D; and there is likely more information leakage. The results are averaged over 50 independent runs; error bars display &#x000B1;2 standard errors.</p>
<p><xref ref-type="fig" rid="F6">Figure 6</xref> shows the results of this experiment. We see that our algorithm&#x00027;s performance remains similar (within each other&#x00027;s the confidence interval) for |<sans-serif>val</sans-serif>(<italic>C</italic><sub>0</sub>)|/|<sans-serif>val</sans-serif>(<italic>C</italic><sub>1</sub>)| &#x02208; {0.25, 0.375}, but significantly worsens when |<sans-serif>val</sans-serif>(<italic>C</italic><sub>0</sub>)|/|<sans-serif>val</sans-serif>(<italic>C</italic><sub>1</sub>)| &#x0003D; 0.5. However, our algorithm continues to perform better than all baselines for all values of |<sans-serif>val</sans-serif>(<italic>C</italic><sub>0</sub>)|/|<sans-serif>val</sans-serif>(<italic>C</italic><sub>1</sub>)|.</p>
<fig id="F6" position="float">
<label>Figure 6</label>
<caption><p>Results of the experiment studying robustness to narrowness of <inline-formula><mml:math id="M148"><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow></mml:math></inline-formula> (Experiment 4, Section 4.5.1). It shows the relation between |<sans-serif>val</sans-serif>(<italic>C</italic><sub>0</sub>)|/|<sans-serif>val</sans-serif>(<italic>C</italic><sub>1</sub>)| and Regret AUC. Our algorithm&#x00027;s performance remains similar (within each other&#x00027;s the confidence interval) for lower values of |<sans-serif>val</sans-serif>(<italic>C</italic><sub>0</sub>)|/|<sans-serif>val</sans-serif>(<italic>C</italic><sub>1</sub>)|, but significantly worsens when |<sans-serif>val</sans-serif>(<italic>C</italic><sub>0</sub>)|/|<sans-serif>val</sans-serif>(<italic>C</italic><sub>1</sub>)| = 0.5. However, our algorithm continues to perform better than all baselines for all values of |<sans-serif>val</sans-serif>(<italic>C</italic><sub>0</sub>)|/|<sans-serif>val</sans-serif>(<italic>C</italic><sub>1</sub>)|.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1346700-g0006.tif"/>
</fig>
</sec>
<sec>
<title>4.5.2 Experiment 5 (size of initial dataset)</title>
<p>The number of samples in the initial dataset <inline-formula><mml:math id="M149"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> would impact the algorithm&#x00027;s resulting policy, for any given <italic>B</italic>. Specifically, we would expect that as the cardinality of <inline-formula><mml:math id="M150"><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub></mml:math></inline-formula> increases, regret reduces. For this experiment, we consider a uniformly exploring logging policy, and vary <inline-formula><mml:math id="M151"><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:math></inline-formula> by setting it to be <inline-formula><mml:math id="M152"><mml:mi>k</mml:mi><mml:mo>&#x000B7;</mml:mo><mml:mo>|</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo><mml:mo>&#x000B7;</mml:mo><mml:mo>|</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:math></inline-formula>, where <italic>k</italic> &#x02208; {0, 0.25, 0.5}. The rest of the setting is the same as in Experiment 2. The results are averaged over 50 independent runs; error bars display &#x000B1;2 standard errors.</p>
<p>The results are shown in <xref ref-type="fig" rid="F7">Figure 7</xref>. We would expect the performance of all algorithms improve with increase in <italic>k</italic> since that would give the agent better starting beliefs; this, indeed, is what we observe. Importantly, our algorithm performs better than all baselines in all these settings. <xref ref-type="fig" rid="F7">Figure 7</xref> broken down by <italic>B</italic> is provided in the <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>.</p>
<fig id="F7" position="float">
<label>Figure 7</label>
<caption><p>Results of the experiment studying robustness to size of the initial dataset (Experiment 5, Section 4.5.2). It shows the relation between <inline-formula><mml:math id="M153"><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo><mml:mo>/</mml:mo><mml:mo>|</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>&#x000D7;</mml:mo><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>|</mml:mo></mml:math></inline-formula> and Regret AUC. As expected, the performance of all algorithms improve with increase in <inline-formula><mml:math id="M154"><mml:mo>|</mml:mo><mml:msub><mml:mrow><mml:mrow><mml:mi mathvariant="script">D</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>L</mml:mi></mml:mrow></mml:msub><mml:mo>|</mml:mo></mml:math></inline-formula>. Importantly, our algorithm performs better than all baselines in all these settings.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1346700-g0007.tif"/>
</fig>
</sec>
<sec>
<title>4.5.3 Experiment 6 (choice of &#x003B2;)</title>
<p>Though we allow the cost function to be arbitrary, this experiment studies our algorithm&#x00027;s performance under two natural choices of &#x003B2;(.) to test its robustness: (1) a constant cost function; that is, <inline-formula><mml:math id="M155"><mml:mi>&#x003B2;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x0221D;</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:math></inline-formula>, and (2) cost function that is inversely proportional to the likelihood of observing the context naturally (i.e., rarer samples are costlier); that is, <inline-formula><mml:math id="M156"><mml:mi>&#x003B2;</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x0221D;</mml:mo><mml:mfrac><mml:mrow><mml:msub><mml:mrow><mml:mi>N</mml:mi></mml:mrow><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub></mml:mrow><mml:mrow><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:mfrac></mml:math></inline-formula>. The rest of the setting is the same as in Experiment 2. The results are averaged over 50 independent runs; error bars display &#x000B1;2 standard errors.</p>
<p><xref ref-type="fig" rid="F8">Figure 8</xref> shows the results. As expected, the choice of cost function does affect performance of all algorithms. However, our algorithm performs better than all algorithms for both cost function choices.</p>
<fig id="F8" position="float">
<label>Figure 8</label>
<caption><p>Results of the experiment studying robustness of Regret AUC of our algorithm to the choice of &#x003B2; (Experiment 6, Section 4.5.3). Though, as expected, the choice of cost function affects performance of all algorithms, our algorithm performs better than all algorithms for both cost function choices.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1346700-g0008.tif"/>
</fig>
</sec>
<sec>
<title>4.5.4 Experiment 7 (misspecification of <inline-formula><mml:math id="M157"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula>)</title>
<p>In real-world applications, the true underlying causal graph may not always be known. In this experiment, we study the impact of mis-specification of <inline-formula><mml:math id="M158"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula> on the performance of our algorithm. Note that the formalism described in Section 2.1 does not necessitate that the agent knows the true underlying causal graph, but rather only that it knows a causal graph such that &#x02119; factorizes according to it. This means that the graph <inline-formula><mml:math id="M159"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula> that the agent knows might include additional arrows not present in the true underlying graph. Intuitively, using such an imperfect graph would result in worsened performance by our algorithm since there are less overlapping information pathways to exploit.</p>
<p>In Experiment 7a, we study this effect empirically by comparing results of Experiment 2 to the same experiment but with the causal graph having an extra edge: <italic>C</italic><sub>1</sub> &#x02192; <italic>Y</italic>. The results are averaged over 25 independent runs; error bars display &#x000B1;2 standard errors. <xref ref-type="fig" rid="F9">Figure 9A</xref> shows the results of the experiment. As expected, performance of our algorithm degrades when there is imperfect knowledge of the true underlying graph. However, our algorithm continues to perform better than all baselines, while also maintaining a similar difference in regret AUC compared to the baselines. <xref ref-type="fig" rid="F9">Figure 9A</xref> broken down by <italic>B</italic> is provided in the <xref ref-type="supplementary-material" rid="SM1">Supplementary material</xref>.</p>
<fig id="F9" position="float">
<label>Figure 9</label>
<caption><p><bold>(A)</bold> Experiment 7a results (misspecification of <inline-formula><mml:math id="M160"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula> in Experiment 2). Though, as expected, imperfect knowledge of the true underlying graph results in reduced performance of our algorithm, it retains its relative better performance. <bold>(B)</bold> Experiment 7b results (misspecification of <inline-formula><mml:math id="M161"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula> in Experiment 3). All algorithms experience mild degradation in performance; our algorithm retains its relative performance compared to baselines.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1346700-g0009.tif"/>
</fig>
<p>In Experiment 7b, we perform a similar analysis to the real-world inspired experiment presented in Section 4.4. Specifically, we compare the case where the true causal graph is known with the cases where there is a misspecification of the causal relationships between the context variables. To capture this, we add one edge (<italic>C</italic><sub>1</sub> &#x02192; <italic>C</italic><sub>0</sub>) to <inline-formula><mml:math id="M162"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula> and then one more edge (<italic>C</italic><sub>2</sub> &#x02192; <italic>C</italic><sub>1</sub>); we compare the performance under these two settings to the case where the true causal graph is known. The results are averaged over 25 independent runs; error bars display &#x000B1;2 standard errors. The results in <xref ref-type="fig" rid="F9">Figure 9B</xref> shows the results. In this case, the deterioration in performance due to misspecification of <inline-formula><mml:math id="M163"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula> is quite small for all algorithms. However, our algorithm continues to perform better than all baselines; it performs the best when the true graph is known.</p>
</sec>
</sec>
<sec>
<title>4.6 Regret behavior for large values of <italic>B</italic></title>
<p><xref ref-type="fig" rid="F3">Figures 3</xref>, <xref ref-type="fig" rid="F4">4</xref> provided regret behavior for small values of <italic>B</italic>. We are primarily interested in such small-budget behavior since that occurs more commonly in practice; for example, budgets exclusively for experimentation in software teams in often quite low.</p>
<p>However, it is also interesting to look at regret behavior as <italic>B</italic> becomes large. Specifically, we increase <italic>B</italic> large enough that all algorithms converge to optimal (or very close to optimal). We do this for Experiments 1 and 2. <xref ref-type="fig" rid="F10">Figures 10A</xref>, <xref ref-type="fig" rid="F10">B</xref> provide the results. Note that the <xref ref-type="fig" rid="F3">Figures 3</xref>, <xref ref-type="fig" rid="F4">4</xref> just zoom into these plots for small <italic>B</italic> (i.e., <italic>B</italic> between 15 and 30).</p>
<fig id="F10" position="float">
<label>Figure 10</label>
<caption><p><bold>(A)</bold> Results of Experiment 1 (provided in Section 4.2) extended for large values of <italic>B</italic>. <bold>(B)</bold> Results of Experiment 2 (provided in Section 4.3) extended for large values of <italic>B</italic>.</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1346700-g0010.tif"/>
</fig>
<p><sans-serif>PropToValue</sans-serif> is the slowest to converge to the optimal policy in both instances, though it demonstrates better low-budget behavior than <sans-serif>EqualAlloc</sans-serif>. <sans-serif>MaxSum</sans-serif> has the best low budget behavior among the baselines because it maximizes the total number of samples within that low budget; it, as <italic>B</italic> gets larger, <sans-serif>EqualAlloc</sans-serif> catches up (and even outperforms it) as it explores the context-action space better. In both experiments, however, our algorithm converges to an optimal policy faster than all baselines.</p>
</sec>
<sec>
<title>4.7 Intuition for better performance of our algorithm</title>
<p>As discussed in Section 2.2, our algorithm balances the trade off between allocating more samples to context-action pairs that are higher value according to its beliefs and allocating more samples for exploration, while taking into account information leakage due to the causal graph. To understand this in more detail, we consider Setting 1 of Experiment 1, and zoom into the case where <italic>B</italic> &#x0003D; 20. We do 50 independent runs and plot the frequency of choosing samples containing different value of <italic>C</italic><sub>1</sub>. We show this for our algorithm and all baselines (except <sans-serif>EqualAlloc</sans-serif> since it is obvious how it allocates).</p>
<p><xref ref-type="fig" rid="F11">Figure 11</xref> shows the results of this experiment. <sans-serif>MaxSum</sans-serif> allocates lesser number of samples than our algorithm to the two context values (<italic>C</italic><sub>1</sub> &#x02208; {0, 1}) that are high value. <sans-serif>PropToValue</sans-serif> over-allocates to these two context values, resulting in poor exploration of other contexts. Our algorithm, in contrast, allocate relatively more to the high-value contexts, while also maintaining good exploration of other contexts.</p>
<fig id="F11" position="float">
<label>Figure 11</label>
<caption><p>Frequency of choosing or encountering each value of <inline-formula><mml:math id="M164"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula>. Highlighted in teal color are the &#x0201C;high-value&#x0201D; contexts (i.e., contexts for which learning the right actions provides higher expected rewards).</p></caption>
<graphic mimetype="image" mime-subtype="tiff" xlink:href="frai-07-1346700-g0011.tif"/>
</fig>
</sec>
</sec>
<sec id="s5">
<title>5 Fairness results</title>
<p>Fairness is becoming an increasingly important angle to discuss when designing machine learning algorithms. A common way to approach fairness is to ensure some subset of variables (assumed given to the algorithm), called &#x0201C;sensitive variables,&#x0201D; is not discriminated against. Specific formal definitions of this discrimination give rise to different notions of fairness in literature (Grgi&#x00107;-Hla&#x0010D;a et al., <xref ref-type="bibr" rid="B7">2016</xref>; Dwork et al., <xref ref-type="bibr" rid="B6">2012</xref>; Kusner et al., <xref ref-type="bibr" rid="B12">2017</xref>; Zuo et al., <xref ref-type="bibr" rid="B34">2022</xref>; Castelnovo et al., <xref ref-type="bibr" rid="B4">2022</xref>).</p>
<sec>
<title>5.1 Counterfactual fairness</title>
<p>Counterfactual fairness is a commonly used notion of individual fairness. Intuitively, a <italic>counterfactually fair</italic> mapping from contexts to actions ensures that the actions mapped to an individual (given by a specific choice of values for the context variables) are the same in a counterfactual world where a subset <inline-formula><mml:math id="M165"><mml:mi>W</mml:mi><mml:mo>&#x02286;</mml:mo><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:math></inline-formula> of sensitive contexts is changed.</p>
<p>For example, if a loan decision policy is counterfactually fair with respect to gender, then a a male applicant would have had the same probability of being granted the loan <italic>had he been</italic> a different gender. Counterfactual fairness requires the knowledge of the causal graph to be able to establish, making it difficult to use in practice; however, when causal graphs are available, it is a powerful notion of individual-level fairness.</p>
<p>In our case, counterfactual fairness can be achieved by setting <inline-formula><mml:math id="M166"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> to contain all the sensitive attributes; that is, by letting <inline-formula><mml:math id="M167"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02287;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula>.</p>
<sec>
<title>5.1.1 Proof of counterfactual fairness</title>
<p>To prove counterfactual fairness, first note that the learned policy is a map <inline-formula><mml:math id="M168"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mtext>&#x000A0;</mml:mtext><mml:mo>:</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x02192;</mml:mo><mml:mstyle mathvariant="sans-serif"><mml:mtext>val</mml:mtext></mml:mstyle><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>X</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula>; during inference, for any given <bold>c</bold><sup><italic>A</italic></sup>, the value of <italic>X</italic> is intervened to be set to <inline-formula><mml:math id="M169"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula>. Following the notation in Pearl (<xref ref-type="bibr" rid="B17">2009a</xref>), we let <inline-formula><mml:math id="M170"><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02190;</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula> denote <inline-formula><mml:math id="M171"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula> in the counterfactual world where the variables in <inline-formula><mml:math id="M172"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> are set equal to <bold>c</bold><sup><italic>B&#x02032;</italic></sup>. To achieve counterfactual fairness [whose definition we draw from the definitions in Kusner et al. (<xref ref-type="bibr" rid="B12">2017</xref>) and Zuo et al. (<xref ref-type="bibr" rid="B34">2022</xref>)], it is sufficient that, for all <bold>c</bold><sup><italic>A</italic></sup>, <bold>c</bold><sup><italic>B</italic></sup>, <bold>c</bold><sup><italic>B&#x02032;</italic></sup>, <italic>x</italic>,</p>
<disp-formula id="E31"><mml:math id="M173"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02190;</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02190;</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>Now, under the assumptions in Section 2.1.4, the conditional independences in <inline-formula><mml:math id="M174"><mml:mrow><mml:mi mathvariant="script">G</mml:mi></mml:mrow></mml:math></inline-formula> imply that we have</p>
<disp-formula id="E32"><mml:math id="M175"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x02200;</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>This gives us that</p>
<disp-formula id="E33"><mml:math id="M176"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msub><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02190;</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mo>&#x02200;</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mi>x</mml:mi></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>which satisfies the counterfactual fairness condition.</p>
</sec>
</sec>
<sec>
<title>5.2 Demographic parity</title>
<p>A common criterion for group-level fairness is Demographic Parity (Kusner et al., <xref ref-type="bibr" rid="B12">2017</xref>). Demographic Parity (DP) requires that the distribution over actions remains the same irrespective of the value of the sensitive variables.</p>
<p>Demographic parity is a popular notion of fairness as it aligns well with most people&#x00027;s understanding of fairness and does not require strong assumptions such as the knowledge of the underlying causal graph to compute. As an example, if a loan decision policy has demographic parity with respect to gender, then the probability of granting a loan given a random male is the same as that for a random individual from any other gender group.</p>
<p>In our case, note that</p>
<disp-formula id="E34"><mml:math id="M177"><mml:mrow><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow></mml:math></disp-formula>
<p>Formally, DP requires that</p>
<disp-formula id="E35"><mml:math id="M178"><mml:mrow><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:mtext>&#x000A0;</mml:mtext><mml:mo>&#x02200;</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:math></disp-formula>
<p>Our algorithm, however, does <italic>not</italic> guarantee DP even if <inline-formula><mml:math id="M179"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02287;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula>:</p>
<disp-formula id="E36"><mml:math id="M180"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mo>=</mml:mo><mml:mi>x</mml:mi><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup><mml:mo>|</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mi>&#x02119;</mml:mi><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mover accent="true"><mml:mrow><mml:mi>&#x003D5;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>=</mml:mo><mml:mi>x</mml:mi></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>which may not equal &#x02119;[&#x003D5; &#x0003D; <italic>x</italic>|<bold>c</bold><sup><italic>B&#x02032;</italic></sup>] since &#x02119;[<bold>c</bold><sup><italic>A</italic></sup>|<bold>c</bold><sup><italic>B</italic></sup>] may not equal &#x02119;[<bold>c</bold><sup><italic>A</italic></sup>|<bold>c</bold><sup><italic>B&#x02032;</italic></sup>]. And since <inline-formula><mml:math id="M181"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>B</mml:mi></mml:mrow></mml:msup><mml:mo>&#x02287;</mml:mo><mml:mi>W</mml:mi></mml:math></inline-formula>, we cannot guarantee demographic parity.</p>
<p>However, we discuss one way through which DP can be achieved, but with a reduction in agent&#x00027;s performance. Specifically, we can achieve DP by ensuring that the agent acts according to a fixed policy irrespective of the value of <bold>c</bold><sup><italic>A</italic></sup>. Intuitively, we construct a fixed policy that maximizes rewards given the agent&#x00027;s learned beliefs.</p>
<p>Specifically, let <inline-formula><mml:math id="M182"><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>&#x0225C;</mml:mo><mml:mover accent="true"><mml:mrow><mml:mo>&#x1D53C;</mml:mo></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mi>Y</mml:mi><mml:mo>|</mml:mo><mml:mi>d</mml:mi><mml:mi>o</mml:mi><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:math></inline-formula>. Assume the fixed policy is probabilistic. Therefore, we&#x00027;re interested in a policy <italic>q</italic> which is a distribution over |<sans-serif>val</sans-serif>(<italic>X</italic>)|. Denoting <italic>q</italic><sup>(<italic>x</italic>)</sup> &#x0225C; <italic>q</italic>(<italic>x</italic>), we solve the following optimization problem:</p>
<disp-formula id="E37"><mml:math id="M183"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mo>&#x0003C;</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>&#x0003E;</mml:mo><mml:mo>=</mml:mo></mml:mtd></mml:mtr><mml:mtr><mml:mtd><mml:mtext>&#x02003;</mml:mtext><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mo>&#x0003C;</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>.</mml:mo><mml:mo>&#x0003E;</mml:mo></mml:mrow></mml:msub><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msup><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mrow><mml:mo>]</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>subject to</p>
<disp-formula id="E38"><mml:math id="M184"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:mi>x</mml:mi></mml:mrow></mml:munder></mml:mstyle><mml:msup><mml:mrow><mml:mi>q</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msup><mml:mo>=</mml:mo><mml:mn>1</mml:mn></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>It is easy to see that one global optimum to this involves assigning a probability of 1 to an action <italic>x</italic> that results in the largest value of <inline-formula><mml:math id="M185"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:mi>x</mml:mi><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:math></inline-formula>. That is, choose <italic>x</italic> such that</p>
<disp-formula id="E39"><mml:math id="M186"><mml:mtable columnalign="left"><mml:mtr><mml:mtd><mml:mi>x</mml:mi><mml:mo>=</mml:mo><mml:mi>a</mml:mi><mml:mi>r</mml:mi><mml:mi>g</mml:mi><mml:mi>m</mml:mi><mml:mi>a</mml:mi><mml:msub><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:msub><mml:mstyle displaystyle="true"><mml:munder class="msub"><mml:mrow><mml:mo>&#x02211;</mml:mo></mml:mrow><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow></mml:munder></mml:mstyle><mml:mover accent="true"><mml:mrow><mml:mi>&#x02119;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo>[</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mstyle mathvariant='bold'><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo>]</mml:mo></mml:mrow><mml:mo>&#x000B7;</mml:mo><mml:mover accent="true"><mml:mrow><mml:mi>&#x003C8;</mml:mi></mml:mrow><mml:mo>^</mml:mo></mml:mover><mml:mrow><mml:mo stretchy="true">(</mml:mo><mml:mrow><mml:msup><mml:mrow><mml:mi>x</mml:mi></mml:mrow><mml:mrow><mml:mi>&#x02032;</mml:mi></mml:mrow></mml:msup><mml:mo>,</mml:mo><mml:msup><mml:mrow><mml:mstyle mathvariant="bold"><mml:mtext>c</mml:mtext></mml:mstyle></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:mrow><mml:mo stretchy="true">)</mml:mo></mml:mrow></mml:mtd></mml:mtr></mml:mtable></mml:math></disp-formula>
<p>and let <italic>q</italic>(<italic>x</italic>) &#x0003D; 1, and <italic>q</italic>(<italic>x</italic>&#x02032;) &#x0003D; 0, &#x02200;<italic>x</italic>&#x02032; &#x02260; <italic>x</italic>. Note that this fixed policy would perform worser on expectation than the context-specific policy learned by the agent in the main part of the paper. This, however, is a cost that can be paid to achieve DP.</p>
<p>As an example of how this might be useful in practice, note that a recruitment agency might be more concerned (perhaps due to legal requirements) about ensuring that it does not discriminate against particular races, genders, etc., even if that leads to suboptimal allocations of jobs. The above approach ensures that such an entity can achieve demographic parity, even though it costs in terms of performance.</p>
<p>As a toy example, suppose there are just two contexts {0, 1} and two actions {0, 1} such that for context 0, the rewards for actions 0 and 1 are 10 and 0, respectively; and for context 1, the corresponding rewards are 0 and 9. Both contexts are equally likely. In this case, the above approach will choose action 0 always and achieve DP with respect to the context variable, but will receive 0 reward for context 1 which is far from optimal. We do not provide a full-fledged characterization of the performance-fairness tradeoffs and reserve that for future work.</p>
</sec>
</sec>
<sec id="s6">
<title>6 Discussion and conclusion</title>
<p>Though exploitation of causal side information in multi-armed bandits has been relatively well-studied, its integration in contextual bandit settings remains much less investigated. This work presented only the second work in this area. Specifically, this paper proposed a new contextual bandit problem formalism where the agent, which has access to qualitative causal side information, can also actively obtain a table of experimental data in one shot, but at a cost and within a budget.</p>
<p>Further, most contextual bandit problems have been studied in the case when contexts are received from the environment. However, there are several real world settings such as marketing campaigns where targeted experiments are possible. This work is one of the very few works to study this setting.</p>
<p>We proposed a novel algorithm based on a new measure similar to entropy, and showed extensive empirical analysis of our algorithm&#x00027;s performance. We also showed theoretical results on soundness and regret. As demonstrated, smartly exploiting information leakage from the causal graph can yield significantly improved performance. Further, it is also possible to achieve certain notions of individual and group fairness, though it might come at a cost of reduced performance.</p>
<p>This work opens up various directions of future research. Fairness of machine learning models is one of the fast-growing areas of research due to the increased interest in responsible development of AI; it is worthwhile to design causal contextual bandit algorithms that meet population-level fairness criteria with minimal impact on performance. It is also useful to investigate more general approaches to fairness such as optimizing under general fairness constraints. Another useful direction is to allow the presence of unobserved confounders in the causal model; while the lack of unmeasured confounders is frequently assumed in causal bandits literature, removing that assumption can provide a qualitative improvement to the setting and make it more widely applicable. It is also interesting to investigate infinite action or context spaces (for example, domains that are continuous) as they often show up in real world scenarios. Another direction is to study contextual bandit algorithms that combine causal discovery with regret optimization can further expand the scope of application by entirely removing the need for a causal graph to be provided. Finally, it might be fruitful to investigate ways to combine one-shot algorithms with sequential algorithms to provide a more comprehensive approach, and to also look at evolving causal graphs.</p>
</sec>
</body>
<back>
<sec sec-type="data-availability" id="s7">
<title>Data availability statement</title>
<p>Publicly available datasets were analyzed in this study. This data can be found at: Zenodo, <ext-link ext-link-type="uri" xlink:href="https://zenodo.org/">https://zenodo.org/</ext-link>, doi: <ext-link ext-link-type="uri" xlink:href="https://doi.org/10.5281/zenodo.5540348">10.5281/zenodo.5540348</ext-link>.</p>
</sec>
<sec sec-type="author-contributions" id="s8">
<title>Author contributions</title>
<p>CS: Writing &#x02013; original draft, Conceptualization, Data curation, Investigation, Methodology, Software, Writing &#x02013; review &#x00026; editing. BR: Writing &#x02013; review &#x00026; editing, Conceptualization, Funding acquisition, Methodology, Supervision, Validation.</p>
</sec>
<sec sec-type="funding-information" id="s9">
<title>Funding</title>
<p>The author(s) declare that no financial support was received for the research, authorship, and/or publication of this article.</p>
</sec>
<ack><p>This paper is based on two chapters from the thesis (Subramanian, <xref ref-type="bibr" rid="B26">2024</xref>).</p>
</ack>
<sec sec-type="COI-statement" id="conf1">
<title>Conflict of interest</title>
<p>The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest.</p>
<p>The author(s) declared that they were an editorial board member of Frontiers, at the time of submission. This had no impact on the peer review process and the final decision.</p>
</sec>
<sec sec-type="disclaimer" id="s10">
<title>Publisher&#x00027;s note</title>
<p>All claims expressed in this article are solely those of the authors and do not necessarily represent those of their affiliated organizations, or those of the publisher, the editors and the reviewers. Any product that may be evaluated in this article, or claim that may be made by its manufacturer, is not guaranteed or endorsed by the publisher.</p>
</sec>
<sec sec-type="supplementary-material" id="s11">
<title>Supplementary material</title>
<p>The Supplementary Material for this article can be found online at: <ext-link ext-link-type="uri" xlink:href="https://www.frontiersin.org/articles/10.3389/frai.2024.1346700/full#supplementary-material">https://www.frontiersin.org/articles/10.3389/frai.2024.1346700/full#supplementary-material</ext-link></p>
<supplementary-material xlink:href="Data_Sheet_1.pdf" id="SM1" mimetype="application/pdf" xmlns:xlink="http://www.w3.org/1999/xlink"/>
</sec>
<fn-group>
<fn id="fn0001"><p><sup>1</sup>Actions in a contextual bandit setting can also be interpreted as Pearl <italic>do</italic>() interventions (Zhang and Bareinboim, <xref ref-type="bibr" rid="B32">2017</xref>; Lattimore et al., <xref ref-type="bibr" rid="B13">2016</xref>); so we use the term &#x0201C;action&#x0201D; and &#x0201C;intervention&#x0201D; interchangeably.</p></fn>
<fn id="fn0002"><p><sup>2</sup>Similar results for multi-armed bandits for best arm identification problems are shown in works such as Lattimore et al. (<xref ref-type="bibr" rid="B13">2016</xref>).</p></fn>
<fn id="fn0003"><p><sup>3</sup>See: <ext-link ext-link-type="uri" xlink:href="https://firebase.google.com/docs/ab-testing">https://firebase.google.com/docs/ab-testing</ext-link> for an example.</p></fn>
<fn id="fn0004"><p><sup>4</sup>See: <ext-link ext-link-type="uri" xlink:href="https://www.persado.com/articles/the-power-of-experimental-design-to-deliver-marketing-insights/">https://www.persado.com/articles/the-power-of-experimental-design-to-deliver-marketing-insights/</ext-link> for an example.</p></fn>
<fn id="fn0005"><p><sup>5</sup>For example, see: <ext-link ext-link-type="uri" xlink:href="https://support.google.com/google-ads/answer/1704368">https://support.google.com/google-ads/answer/1704368</ext-link></p></fn>
<fn id="fn0006"><p><sup>6</sup>For example, see: <ext-link ext-link-type="uri" xlink:href="https://scale.com/blog/netflix-recommendation-personalization">https://scale.com/blog/netflix-recommendation-personalization</ext-link></p></fn>
<fn id="fn0007"><p><sup>7</sup>By qualitative, we mean that the agent can know the causal <italic>graph</italic>, but not the conditional probability distributions of the variables. See Section 2.1 for a more detailed discussion.</p></fn>
<fn id="fn0008"><p><sup>8</sup>In contrast, in a standard contextual bandit setting, the agent would repeatedly observe <bold>c</bold><sup><italic>A</italic></sup>, respond with an <italic>x</italic>, and observe outcomes.</p></fn>
<fn id="fn0009"><p><sup>9</sup><italic>do</italic>(<italic>X</italic> &#x0003D; <italic>x</italic>) is a standard operation in causal inference that models interventions, where parents of <italic>X</italic> are removed, and <italic>X</italic> is set equal to <italic>x</italic>. See Pearl (<xref ref-type="bibr" rid="B18">2009b</xref>, <xref ref-type="bibr" rid="B19">2019</xref>) for more discussion on the <italic>do</italic>() operation. <italic>do</italic>(<italic>X</italic> &#x0003D; <italic>x</italic>) is succinctly written <italic>do</italic>(<italic>x</italic>).</p></fn>
<fn id="fn0010"><p><sup>10</sup>That is, if <inline-formula><mml:math id="M49"><mml:msup><mml:mrow><mml:mrow><mml:mi mathvariant="script">C</mml:mi></mml:mrow></mml:mrow><mml:mrow><mml:mi>A</mml:mi></mml:mrow></mml:msup></mml:math></inline-formula> contains all its ancestors.</p></fn>
<fn id="fn0011"><p><sup>11</sup>&#x0201C;Belief&#x0201D; is a commonly used term to mean the posterior distribution of the parameters of the distribution of interest (Russo et al., <xref ref-type="bibr" rid="B21">2017</xref>).</p></fn>
<fn id="fn0012"><p><sup>12</sup>Information leakage arises from shared pathways in <inline-formula><mml:math id="M65"><mml:mrow><mml:mi mathvariant="script">M</mml:mi></mml:mrow></mml:math></inline-formula>; or equivalently, due to shared CPDs in the factorization of &#x02119;.</p></fn>
<fn id="fn0013"><p><sup>13</sup><ext-link ext-link-type="uri" xlink:href="https://docs.scipy.org/doc/scipy/reference/generated/scipy.optimize.differential_evolution.html">https://docs.scipy.org/doc/scipy/reference/generated/scipy.optimize.differential_evolution.html</ext-link></p></fn>
</fn-group>
<ref-list>
<title>References</title>
<ref id="B1">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Agrawal</surname> <given-names>S.</given-names></name> <name><surname>Goyal</surname> <given-names>N.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Analysis of Thompson sampling for the multi-armed Bandit problem,&#x0201D;</article-title> in <italic>Proceedings of the 25th Annual Conference on Learning Theory, Vol. 23 of PMLR</italic> (Edinburgh), <fpage>39.1</fpage>&#x02013;<lpage>39.26</lpage>.</citation>
</ref>
<ref id="B2">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ameko</surname> <given-names>M. K.</given-names></name> <name><surname>Beltzer</surname> <given-names>M. L.</given-names></name> <name><surname>Cai</surname> <given-names>L.</given-names></name> <name><surname>Boukhechba</surname> <given-names>M.</given-names></name> <name><surname>Teachman</surname> <given-names>B. A.</given-names></name> <name><surname>Barnes</surname> <given-names>L. E.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Offline contextual multi-armed bandits for mobile health interventions: a case study on emotion regulation,&#x0201D;</article-title> in <italic>Proceedings of the 14th ACM Conference on Recommender Systems</italic> (New York, NY), <fpage>249</fpage>&#x02013;<lpage>258</lpage>.</citation>
</ref>
<ref id="B3">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Bouneffouf</surname> <given-names>D.</given-names></name> <name><surname>Rish</surname> <given-names>I.</given-names></name> <name><surname>Aggarwal</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Survey on applications of multi-armed and contextual bandits,&#x0201D;</article-title> in <italic>2020 IEEE Congress on Evolutionary Computation (CEC)</italic> (Glasgow), <fpage>1</fpage>&#x02013;<lpage>8</lpage>.</citation>
</ref>
<ref id="B4">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Castelnovo</surname> <given-names>A.</given-names></name> <name><surname>Crupi</surname> <given-names>R.</given-names></name> <name><surname>Greco</surname> <given-names>G.</given-names></name> <name><surname>Regoli</surname> <given-names>D.</given-names></name> <name><surname>Penco</surname> <given-names>I. G.</given-names></name> <name><surname>Cosentini</surname> <given-names>A. C.</given-names></name></person-group> (<year>2022</year>). <article-title>A clarification of the nuances in the fairness metrics landscape</article-title>. <source>Sci. Rep</source>. <volume>12</volume>:<fpage>4209</fpage>. <pub-id pub-id-type="doi">10.1038/s41598-022-07939-1</pub-id><pub-id pub-id-type="pmid">35273279</pub-id></citation></ref>
<ref id="B5">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dulac-Arnold</surname> <given-names>G.</given-names></name> <name><surname>Levine</surname> <given-names>N.</given-names></name> <name><surname>Mankowitz</surname> <given-names>D. J.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Paduraru</surname> <given-names>C.</given-names></name> <name><surname>Gowal</surname> <given-names>S.</given-names></name> <etal/></person-group>. (<year>2021</year>). <article-title>Challenges of real-world reinforcement learning: definitions, benchmarks and analysis</article-title>. <source>Machine Learn</source>. <volume>110</volume>, <fpage>2419</fpage>&#x02013;<lpage>2468</lpage>. <pub-id pub-id-type="doi">10.1007/s10994-021-05961-4</pub-id></citation>
</ref>
<ref id="B6">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Dwork</surname> <given-names>C.</given-names></name> <name><surname>Hardt</surname> <given-names>M.</given-names></name> <name><surname>Pitassi</surname> <given-names>T.</given-names></name> <name><surname>Reingold</surname> <given-names>O.</given-names></name> <name><surname>Zemel</surname> <given-names>R.</given-names></name></person-group> (<year>2012</year>). <article-title>&#x0201C;Fairness through awareness,&#x0201D;</article-title> in <italic>Proceedings of the 3rd Innovations in Theoretical Computer Science Conference</italic> (New York, NY), <fpage>214</fpage>&#x02013;<lpage>226</lpage>.</citation>
</ref>
<ref id="B7">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Grgi&#x00107;-Hla&#x0010D;a</surname> <given-names>N.</given-names></name> <name><surname>Zafar</surname> <given-names>M. B.</given-names></name> <name><surname>Gummadi</surname> <given-names>K. P.</given-names></name> <name><surname>Weller</surname> <given-names>A.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;The case for process fairness in learning: feature selection for fair decision making,&#x0201D;</article-title> in <italic>Symposium on Machine Learning and the Law at the 29th Conference on Neural Information Processing Systems</italic>. Barcelona.</citation>
</ref>
<ref id="B8">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Guo</surname> <given-names>R.</given-names></name> <name><surname>Cheng</surname> <given-names>L.</given-names></name> <name><surname>Li</surname> <given-names>J.</given-names></name> <name><surname>Hahn</surname> <given-names>P. R.</given-names></name> <name><surname>Liu</surname> <given-names>H.</given-names></name></person-group> (<year>2020</year>). <article-title>A survey of learning causality with data: problems and methods</article-title>. <source>ACM Comput. Surv</source>. <volume>53</volume>, <fpage>1</fpage>&#x02013;<lpage>37</lpage>. <pub-id pub-id-type="doi">10.1145/3397269</pub-id></citation>
</ref>
<ref id="B9">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Han</surname> <given-names>Y.</given-names></name> <name><surname>Zhou</surname> <given-names>Z.</given-names></name> <name><surname>Zhou</surname> <given-names>Z.</given-names></name> <name><surname>Blanchet</surname> <given-names>J.</given-names></name> <name><surname>Glynn</surname> <given-names>P. W.</given-names></name> <name><surname>Ye</surname> <given-names>Y.</given-names></name></person-group> (<year>2020</year>). <article-title>Sequential batch learning in finite-action linear contextual bandits</article-title>. <source>arXiv [preprint]</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2004.06321</pub-id></citation>
</ref>
<ref id="B10">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Joachims</surname> <given-names>T.</given-names></name> <name><surname>Swaminathan</surname> <given-names>A.</given-names></name> <name><surname>Rijke</surname> <given-names>M. d. R.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Deep learning with logged bandit feedback,&#x0201D;</article-title> in <italic>Proceedings of the Sixth International Conference on Learning Representations</italic>. Vancouver, BC.</citation>
</ref>
<ref id="B11">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Koller</surname> <given-names>D.</given-names></name> <name><surname>Friedman</surname> <given-names>N.</given-names></name></person-group> (<year>2009</year>). <source>Probabilistic Graphical Models: Principles and Techniques</source>. <publisher-loc>Cambridge, MA</publisher-loc>: <publisher-name>The MIT Press</publisher-name>.</citation>
</ref>
<ref id="B12">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Kusner</surname> <given-names>M. J.</given-names></name> <name><surname>Loftus</surname> <given-names>J.</given-names></name> <name><surname>Russell</surname> <given-names>C.</given-names></name> <name><surname>Silva</surname> <given-names>R.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Counterfactual fairness,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems, Vol. 30</source>, eds. I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (New York, NY: CA), <fpage>4069</fpage>&#x02013;<lpage>4079</lpage>.</citation>
</ref>
<ref id="B13">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lattimore</surname> <given-names>F.</given-names></name> <name><surname>Lattimore</surname> <given-names>T.</given-names></name> <name><surname>Reid</surname> <given-names>M. D.</given-names></name></person-group> (<year>2016</year>). <article-title>&#x0201C;Causal bandits: learning good interventions via causal inference,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems 29, Vol. 29</source>, eds. D. Lee, M. Sugiyama, U. Luxburg, I. Guyon, and R. Garnett (Barcelona), <fpage>1189</fpage>&#x02013;<lpage>1197</lpage>.</citation>
</ref>
<ref id="B14">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Lattimore</surname> <given-names>T.</given-names></name> <name><surname>Szepesv&#x000E1;ri</surname> <given-names>C.</given-names></name></person-group> (<year>2020</year>). <source>Bandit Algorithms</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation>
</ref>
<ref id="B15">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Liu</surname> <given-names>B.</given-names></name> <name><surname>Wei</surname> <given-names>Y.</given-names></name> <name><surname>Zhang</surname> <given-names>Y.</given-names></name> <name><surname>Yan</surname> <given-names>Z.</given-names></name> <name><surname>Yang</surname> <given-names>Q.</given-names></name></person-group> (<year>2018</year>). <article-title>&#x0201C;Transferable contextual bandit for cross-domain recommendation,&#x0201D;</article-title> in <italic>Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 32</italic> (New York, NY: LA), <fpage>3619</fpage>&#x02013;<lpage>3626</lpage>.</citation>
</ref>
<ref id="B16">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Lu</surname> <given-names>Y.</given-names></name> <name><surname>Meisami</surname> <given-names>A.</given-names></name> <name><surname>Tewari</surname> <given-names>A.</given-names></name> <name><surname>Yan</surname> <given-names>W.</given-names></name></person-group> (<year>2020</year>). <article-title>&#x0201C;Regret analysis of bandit problems with causal background knowledge,&#x0201D;</article-title> in <italic>Proceedings of the 36th Conference on Uncertainty in Artificial Intelligence (UAI), Vol. 124 of PMLR</italic> (New York, NY: Virtual Event), <fpage>141</fpage>&#x02013;<lpage>150</lpage>.</citation>
</ref>
<ref id="B17">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pearl</surname> <given-names>J.</given-names></name></person-group> (<year>2009a</year>). <article-title>Causal inference in statistics: an overview</article-title>. <source>Stat. Surv</source>. <volume>3</volume>, <fpage>96</fpage>&#x02013;<lpage>146</lpage>. <pub-id pub-id-type="doi">10.1214/09-SS057</pub-id></citation>
</ref>
<ref id="B18">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Pearl</surname> <given-names>J.</given-names></name></person-group> (<year>2009b</year>). <source>Causality, 2nd Edn</source>. <publisher-loc>Cambridge</publisher-loc>: <publisher-name>Cambridge University Press</publisher-name>.</citation>
</ref>
<ref id="B19">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Pearl</surname> <given-names>J.</given-names></name></person-group> (<year>2019</year>). <article-title>On the Interpretation of do(x)</article-title>. <source>J. Causal Infer</source>. <volume>7</volume>:<fpage>2002</fpage>. <pub-id pub-id-type="doi">10.1515/jci-2019-2002</pub-id><pub-id pub-id-type="pmid">27887028</pub-id></citation></ref>
<ref id="B20">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Ren</surname> <given-names>Z.</given-names></name> <name><surname>Zhou</surname> <given-names>Z.</given-names></name> <name><surname>Kalagnanam</surname> <given-names>J. R.</given-names></name></person-group> (<year>2022</year>). <article-title>Batched learning in generalized linear contextual bandits with general decision sets</article-title>. <source>IEEE Contr. Syst. Lett</source>. <volume>6</volume>, <fpage>37</fpage>&#x02013;<lpage>42</lpage>. <pub-id pub-id-type="doi">10.1109/LCSYS.2020.3047601</pub-id></citation>
</ref>
<ref id="B21">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Russo</surname> <given-names>D.</given-names></name> <name><surname>Roy</surname> <given-names>B. V.</given-names></name> <name><surname>Kazerouni</surname> <given-names>A.</given-names></name> <name><surname>Osband</surname> <given-names>I.</given-names></name></person-group> (<year>2017</year>). <article-title>A tutorial on thompson sampling</article-title>. <source>arXiv [preprint]</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1707.02038</pub-id></citation>
</ref>
<ref id="B22">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sawant</surname> <given-names>N.</given-names></name> <name><surname>Namballa</surname> <given-names>C. B.</given-names></name> <name><surname>Sadagopan</surname> <given-names>N.</given-names></name> <name><surname>Nassif</surname> <given-names>H.</given-names></name></person-group> (<year>2018</year>). <article-title>Contextual multi-armed bandits for causal marketing</article-title>. <source>arXiv [preprint]</source>. <pub-id pub-id-type="doi">10.48550/arXiv.1810.01859</pub-id></citation>
</ref>
<ref id="B23">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Sen</surname> <given-names>R.</given-names></name> <name><surname>Shanmugam</surname> <given-names>K.</given-names></name> <name><surname>Dimakis</surname> <given-names>A. G.</given-names></name> <name><surname>Shakkottai</surname> <given-names>S.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Identifying best interventions through online importance sampling,&#x0201D;</article-title> in <italic>Proceedings of the 34th International Conference on Machine Learning, Vol. 70 of PMLR</italic> (New York, NY), <fpage>3057</fpage>&#x02013;<lpage>3066</lpage>.</citation>
</ref>
<ref id="B24">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Settles</surname> <given-names>B.</given-names></name></person-group> (<year>2012</year>). <source>Active Learning, 1st Edn</source>. <publisher-loc>Cham</publisher-loc>: <publisher-name>Springer</publisher-name>.</citation>
</ref>
<ref id="B25">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Storn</surname> <given-names>R.</given-names></name> <name><surname>Price</surname> <given-names>K.</given-names></name></person-group> (<year>1997</year>). <article-title>Differential evolution&#x02014;a simple and efficient heuristic for global optimization over continuous spaces</article-title>. <source>J. Glob. Optimizat</source>. <volume>11</volume>, <fpage>341</fpage>&#x02013;<lpage>359</lpage>. <pub-id pub-id-type="doi">10.1023/A:1008202821328</pub-id></citation>
</ref>
<ref id="B26">
<citation citation-type="book"><person-group person-group-type="author"><name><surname>Subramanian</surname> <given-names>C.</given-names></name></person-group> (<year>2024</year>). <source>Causal Contextual Bandits</source> (<publisher-loc>Ph. D. thesis</publisher-loc>). Indian Institute of Technology Madras, Chennai, India.</citation>
</ref>
<ref id="B27">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Subramanian</surname> <given-names>C.</given-names></name> <name><surname>Ravindran</surname> <given-names>B.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Causal contextual bandits with targeted interventions,&#x0201D;</article-title> in <italic>Proceedings of the Tenth International Conference on Learning Representations (ICLR 2022)</italic>. Appleton, WI: Virtual Event.</citation>
</ref>
<ref id="B28">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Swaminathan</surname> <given-names>A.</given-names></name> <name><surname>Joachims</surname> <given-names>T.</given-names></name></person-group> (<year>2015a</year>). <article-title>Batch learning from logged bandit feedback through counterfactual risk minimization</article-title>. <source>J. Machine Learn. Res</source>. <volume>16</volume>, <fpage>1731</fpage>&#x02013;<lpage>1755</lpage>. <pub-id pub-id-type="doi">10.5555/2789272.2886805</pub-id></citation>
</ref>
<ref id="B29">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Swaminathan</surname> <given-names>A.</given-names></name> <name><surname>Joachims</surname> <given-names>T.</given-names></name></person-group> (<year>2015b</year>). <article-title>&#x0201C;Counterfactual risk minimization: learning from logged bandit feedback,&#x0201D;</article-title> in <italic>Proceedings of the 32nd International Conference on International Conference on Machine Learning, Vol. 37</italic> (Lille), <fpage>814</fpage>&#x02013;<lpage>823</lpage>.</citation>
</ref>
<ref id="B30">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Wu</surname> <given-names>H.</given-names></name> <name><surname>Srikant</surname> <given-names>R.</given-names></name> <name><surname>Liu</surname> <given-names>X.</given-names></name> <name><surname>Jiang</surname> <given-names>C.</given-names></name></person-group> (<year>2015</year>). <article-title>&#x0201C;Algorithms with logarithmic or sublinear regret for constrained contextual bandits,&#x0201D;</article-title> in <italic>Advances in Neural Information Processing Systems 28</italic> (Montreal, QC), <fpage>433</fpage>&#x02013;<lpage>441</lpage>.</citation>
</ref>
<ref id="B31">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Yabe</surname> <given-names>A.</given-names></name> <name><surname>Hatano</surname> <given-names>D.</given-names></name> <name><surname>Sumita</surname> <given-names>H.</given-names></name> <name><surname>Ito</surname> <given-names>S.</given-names></name> <name><surname>Kakimura</surname> <given-names>N.</given-names></name> <name><surname>Fukunaga</surname> <given-names>T.</given-names></name> <etal/></person-group>. (<year>2018</year>). <article-title>&#x0201C;Causal bandits with propagating inference,&#x0201D;</article-title> in <italic>Proceedings of the 35th International Conference on Machine Learning, Vol. 80 of PMLR</italic> (Stockholm), <fpage>5512</fpage>&#x02013;<lpage>5520</lpage>.</citation>
</ref>
<ref id="B32">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>J.</given-names></name> <name><surname>Bareinboim</surname> <given-names>E.</given-names></name></person-group> (<year>2017</year>). <article-title>&#x0201C;Transfer learning in multi-armed bandits: a causal approach,&#x0201D;</article-title> in <italic>Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence</italic> (Melbourne, VIC), <fpage>1340</fpage>&#x02013;<lpage>1346</lpage>.</citation>
</ref>
<ref id="B33">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zhang</surname> <given-names>Z.</given-names></name> <name><surname>Ji</surname> <given-names>X.</given-names></name> <name><surname>Zhou</surname> <given-names>Y.</given-names></name></person-group> (<year>2022</year>). <article-title>Almost optimal batch-regret tradeoff for batch linear contextual bandits</article-title>. <source>arXiv [preprint]</source>. <pub-id pub-id-type="doi">10.48550/arXiv.2110.08057</pub-id></citation>
</ref>
<ref id="B34">
<citation citation-type="journal"><person-group person-group-type="author"><name><surname>Zuo</surname> <given-names>A.</given-names></name> <name><surname>Wei</surname> <given-names>S.</given-names></name> <name><surname>Liu</surname> <given-names>T.</given-names></name> <name><surname>Han</surname> <given-names>B.</given-names></name> <name><surname>Zhang</surname> <given-names>K.</given-names></name> <name><surname>Gong</surname> <given-names>M.</given-names></name></person-group> (<year>2022</year>). <article-title>&#x0201C;Counterfactual fairness with partially known causal graph,&#x0201D;</article-title> in <source>Advances in Neural Information Processing Systems, Vol. 35</source>, eds. S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (New York, MY: LA), <fpage>1238</fpage>&#x02013;<lpage>1252</lpage>.</citation>
</ref>
</ref-list>
</back>
</article>