A line drawing of a butterfly at the end of a looping dashed flight path.

Teaching models to reason about harm to animals

Humans are starting to use AIs to help make a wide range of decisions, some of which could harm or benefit animals. AI should take animal welfare into account along with other ethical considerations, but current training data does not teach them how.

Research on alignment finetuning⁠1⁠2 shows how it can be taught: the reasons behind aligned behaviors matter just as much as the behaviors themselves.

Two complementary techniques proved especially effective:

  1. Synthetic document finetuning
    Pretraining-style documents from a world where the target model is already aligned to certain behaviors. This reinforces the existence of an aligned persona.
  2. Difficult advice Q&A
    Depictions of an AI assistant advising users on realistic and complex dilemmas. This teaches the application of reasoning across a range of possible scenarios.

Following this research, we built two pipelines that synthesize the missing training data for welfare considerations of animals and other sentient beings.

The outputs show models acting beneficially towards morally relevant animals while still adhering to broader alignment guidelines and respecting user autonomy.

The two datasets

Synthetic documentsDifficult advice Q&A
outputDiverse artifacts from a world where a model already reasons responsibly about animal welfare.AI coaching users through ethical dilemmas involving disenfranchised third parties (e.g. animals).
output formatBlogs, interviews, encyclopedia entries, forum threads.One user dilemma in, one assistant answer out.
what it is forMidtrainingSupervised fine-tuning QA
pipelinePipelinePipeline
example datasetExample datasetExample dataset

Walk through either pipeline