When Good Algorithms Go Wrong: How AI Recommendations Can Promote Worse Behaviour in Matching Markets
Choosing the right school for your child, picking the perfect job, or even finding the ideal community for a newly arrived refugee to settle in are life-altering yet difficult decisions to make. Where people used to choose based on prior experience, advice, or brochures, now they can use artificial intelligence (AI) tools. These tools can take historical data to predict where a student or employee might thrive. On paper, accurate predictions help people make better choices. But do they lead to better outcomes over the years?
Not always, as it turns out. The problem is that institutions can be motivated to change their behaviour in order to influence these predictions, even if such actions can ultimately harm both the institutions and their people. My recent research (together with Yuhao Du, Kenneth Joseph and Anikó Hannák) delved into how and why this can happen, taking high school placements as an example.
Altering predictions via adversarial interactions
Consider a system in which school placements are determined by central planners based on students’ declared preferences. (Similar centralized matching mechanisms are also used in other contexts, such as placing refugees in particular states or cantons according to their predicted likelihood of employment.)
Students may use an AI recommender to identify their favourite options, with the AI looking at the performance of past students at all the schools. If students with a weak mathematics background have done well at School A, the AI will recommend School A to such students the following year. If these students follow the AI’s recommendations, planners will see School A given preference by many students who are weak in maths.
But School A can change predictions by acting differently. Imagine that the school wants to attract more students with a strong mathematics background in future years, perhaps because although it has good programmes to help students with a weaker foundation, these are costly and time-intensive. By withholding a bit of academic support or extracurricular funding from its current batch of students with weaker backgrounds, School A can make the AI evaluate it as less attractive to such students, ultimately generating fewer recommendations to similar students in the next academic cycle. This enables School A to be matched with the student profiles it prefers.
So schools can change their behaviour in order to shift future predictions. This is what we call an adversarial interaction attack.
Higher accuracy means higher risk
When developing AI systems, engineers often aim for high accuracy – meaning that system predictions should be as close as possible to actual observed outcomes. Intuitively, this is helpful. If the AI tells a student that School A will prepare them for a great score on their final exam, and it turns out to be true, that student can make good, informed choices based on the AI’s advice. It also means that the student can trust the recommendations they get.
However, accuracy is measured with respect to what the AI observes. If the system is accurate, any shift in how a school acts will also quickly be learned by the AI. Moreover, if students trust it, the AI recommendations will be reflected in students’ expressed preferences. If the AI is seen as unreliable, no one listens to it anyway, so gaming it is not worth the schools’ effort.
The above explains a result that initially may sound surprising: if AI makes more accurate and trusted predictions, then institutions can have a higher incentive to implement adversarial interaction attacks.
How does this affect fairness?
In our example, this means that schools invest fewer resources in student preparation. But not all students are impacted equally. Our analysis suggests that students who require more assistance would experience worse declines in outcomes. Bearing in mind that their need for assistance may solely result from having had fewer opportunities in the past, AI recommendations can exacerbate inequalities, irrespective of an individual’s potential.
Does this mean that we should give up on outcome-based recommendations and return to the old way of doing things? This is not the only option. The main takeaway is that we should no longer design matching mechanisms (the process of allocating school places to students) and AI systems in isolation. Traditionally, engineers and computer scientists design and analyse the AI recommenders, while economists or administrators handle the matching mechanisms. However, in isolation, we miss the feedback loop in which strategic actions may aim to shift AI predictions to achieve different future matchings.
Alternative solutions exist. For example, we can keep a lower accuracy in our prediction model in order to discourage interaction attacks. Alternatively, we can choose mechanisms that allow institutions to express their true preferences unconstrained and take these preferences into account.
Solving the problem demands a wider perspective
I am now investigating even more interventions and application domains, with specific attention to social equity. In a joint project between ETH Zürich and Eindhoven University of Technology in the Netherlands, working with Kai Zhang, Sophia Lahrech, Florian Dörfler, Giulia de Pasquale and Valentina Breschi, we are looking at the problem of refugee assignment. In this case, the social planner chooses matchings that maximize the integration of refugees, as predicted by AI models. However, our preliminary results suggest that locations being matched to refugees could manipulate the system using adversarial interactions that postpone refugee employment – just like in the school example.
On the intervention side, we have shown that a social planner could lower the incentive for adversarial interactions by limiting how much an AI model can change over time. (This is similar to the solution in the school choice setting that retained some noise in the AI predictions.)
These adversarial interaction attacks are one example of what can go wrong when intricate technological systems are deployed within already complex social contexts. In my research, I aim to bring that social complexity into the calculation to support better system evaluation and design. I use a mathematical language to combine multiple domains in one model: not only the technical systems (developed in Computer Science), but also the individual and collective human behaviour involved (found in Psychology, Sociology and Economics), as well as desired system properties (formulated by Ethics, Philosophy and Law).
By bringing all these perspectives together, we can analyze the sociotechnical system holistically. It is then possible to design systems that not only optimize for short-term efficiency but also account for social equity alongside other long-term desiderata.
So for schools, we extended school-choice models widely used in economics with an AI component, and used game-theoretic analysis to uncover the interaction attacks. We then performed a similar investigation for the refugee setting and relied on methods from optimization and control to propose more robust decision-support systems.
But there is so much more to investigate. As AI systems get integrated into new decision-making processes and evolve, new vulnerabilities emerge. What will happen in the job market as we rely more on automated recommendations for job choice? Will employers have an incentive to attack the system, and, if so, would the same interventions work? These questions remain open for the future. What this work showed us is that we need an integrated multidisciplinary analysis to ensure that AI systems support the healthy development of our society, not only in the near future, but also in the long run.