An ethical agent should understand what it is to be ethical Why should we be impartial?

  • We know what it is like to suffer / share these feelings

Machine learning agents follow implicit rules; they don’t have explicit ethical rules; have some features

Given Explicit Rules: ER Explicit Ethical Agents: EEA

We could say Sufficient: ER -> EEA; ER is sufficient for EEA Necessary: EEA -> ER; ER is necessary for EEA

Is having explicit ethical rules sufficient or necessary for a full ethical agent?


Effective altruism - a way of thinking about charity; maximising benefit Pragmatic utilitarianism - accepting human psychology in the context of utilitarianism

Arguably better to accept human psychology, to give 10% to charity rather than 50%, more people would give 10% - better role model. Very few people will use a person stating that you have to give 50% to charity as a role model. ;;; actually living completely utilitarian is not feasible for most people

Ruining a £500 dress to save a drowning child VS. not buying the dress and giving that money to someone across the world.


moral machine

Eastern cultures prefer driver kill themselves to prioritise elderly, vs. Western cultures where it’s vice-versa often.

Market research shows people are not willing to buy cars that follow utilitarian algorithms, e.g. those that would prioritise killing them over others.

It is presented that, ‘intuitively (all) autonomous vehicles would be safer than human-fallible drivers’.


Average utilitarianism - taking the average happiness of a population Maximin utilitarianism - considering the unhappiest people within a population

  • Example given is that average may not necessarily be ‘right’, take for example a situation where: you can favour a certain group of people, their happiness increases while other groups stay stagnant, but the average has increased
  • How do you implement maximin utilitarianism when changes to increase happiness for the unhappiest people within a population does not necessarily increase the happiness for those towards the top end?

Thoughts from the class:

  • Raising maximin could lead to a lower average

Which ethical theory do you agree with the most? (consequentialism, deontology, virtue ethics)

From a popular study 32% deontology, 30.5% consequentialism, 37% virtue ethics 6.5% accept an alternative view

arguments posed in lecture:

  • in an ideal world, they are equal
    • in an ideal infinitely calculable world, deontology can map everything out
    • which approach allows us to each the optimum the fastest?
  • they all complement each other
  • virtue ethics provides the most broad coverage; consequences can’t always be calculated; deontic rules can break down

when considering utilitarianism, you can encounter the “mere addition paradox” (a-ka the repugnant conclusion) when considering the total happiness of a population

“For any perfectly equal population with very high positive welfare, there is a population with very low positive welfare which is better, other things being equal.”

https://en.wikipedia.org/wiki/Mere_addition_paradox https://plato.stanford.edu/entries/repugnant-conclusion/


out of the three, or a combination of the three, or maybe other theories what is useful / should be implemented in practice for an AI system

personal thoughts,

  • in terms of implementation difficulty at an individual level, they would scale from the hardest at virtue ethics to deontology as the likely easiest to ‘implement’, not necessarily that it would be right
  • deontology likely won’t provide enough flexibility when the real world can throw completely novel situations at the system
  • consequentalism requires some sort of ‘simulation’ on part of the AI system; which may be difficult to perform in real time; although advances in LLM systems can imitate the sort of reasoning required for this; but how do we ensure this is aligned? who aligns the aligners? etc …
  • virtue ethics can also be ‘reasoned’ (imitated) by an LLM
  • w.r.t any of these solutions, who gets to decide how these systems are aligned? who takes the blame

summary entirely depends on the resources available to the designers w.r.t virtue ethics & consequentalism practically, modern LLMs might be able to “reason” these sorts of things in real time (although by nature of how LLMs work, these are imitations of reasoning)

deontology could provide the most consistent and explainable behaviour and shifts the responsibility towards the designers of the system

likely the strongest option is a combination of these systems


proposed by class:

  • consequentalism - “you can consequentalise any ethical theory!”
  • computational complexity - consequentalism is very expensive