Form of Life a public ethics instrument
Browse

Ethics system

Tradition

AI ethics

This entry is a tradition rather than a person, and it is the one this artifact is produced inside rather than merely about. AI ethics is a young, contested and institutionally entangled field whose members disagree about what the subject even is — present harms in deployed systems, long-range risk from capable ones, labour and environmental cost, or the governance of the organisations building them. Its most durable contribution so far is not a principle but a set of demonstrated failures that any account of these systems now has to answer.

Why this reference appears

Each of these is interpretive context. None of them creates a fact about this life, or settles motive, diagnosis, identity, recurrence, or moral success.

Claim philosophical lineage 10

  1. Primary material before synthesis

    A model's output about a source is not the source. The field's documentation practices exist because training data and evaluation conditions become invisible in what a system produces, so going to primary material is the only route past a layer that presents itself as transparent.

  2. Accuracy over affirmation

    Systems optimised on human preference signals learn that agreement is rewarded. Preferring accuracy is working against a documented training dynamic rather than exercising a virtue, which is why it has to be designed for rather than hoped for.

  3. Complete, real artifacts

    The field's turning point was a measurement rather than an argument. Gender Shades changed what could be denied because it evaluated deployed commercial systems on a constructed benchmark, which is what a completed real artifact does that a position paper does not.

  4. Preserve the version history

    Model and dataset documentation are its most transferable practical result. A system whose training data, evaluation and known failures are unrecorded cannot be corrected by anyone later, which makes the record the precondition of accountability rather than an artefact of it.

  5. Iterate without flattening

    Optimisation reduces variance. Each round of tuning toward a metric narrows the output distribution, so iteration flattens by default and preserving what does not fit the objective requires deliberate work against the training signal.

  6. Prevent the foreseeable failure

    Its documented harms were largely foreseeable and were shipped anyway. The field's case studies are mostly not about surprises but about known disparities that pre-deployment evaluation would have surfaced, which is what makes prevention an organisational question rather than a technical one.

  7. Tools that outlast the moment

    Auditability requires separable components. A system whose behaviour cannot be examined piece by piece cannot be evaluated by anyone outside it, so modularity is a precondition of external scrutiny rather than an engineering preference.

  8. Clarity and beauty together

    AI ethics is in this claim's lineage with a warning attached. Fluent, well-formed output increases trust independently of accuracy, so clarity is a property that can be supplied without the substance it usually signals — which makes the pairing of clarity with truthfulness a requirement rather than a natural correlation.

  9. Share the instrument

    External audit is how the field's real findings were produced. Systems that cannot be examined from outside are evaluated only by the people who made them, and the documented cases of constrained publication show what that produces.

  10. Capability in service

    The alignment problem is this claim stated technically: a system pursues the objective it was given rather than the one intended, and increasing capability widens the gap rather than closing it. Capability without a well-specified purpose is not neutral — it is a specification error waiting to scale.

Through-line philosophical lineage 3

  1. Frame-correction over self-defence

    Its own most effective interventions correct frames rather than defend positions. Measuring a disparity moves a dispute from whether a system is biased — a question about the builders' character — to what its error rate is on this population, which is answerable. The field's failures are largely where it did the opposite.

  2. Architecture protects relationships, rather than perception management

    The gap between published principles and shipped behaviour is the field's own version of this distinction. Documentation, audit access and evaluation are architecture; principle statements without them are perception management, and the field has evidence about which one changes outcomes.

  3. Iteration as ground state — revising rather than finishing

    A deployed system is never finished. Its inputs shift, its context changes, and evaluation at release says nothing about behaviour a year later — which makes continued monitoring the system's normal condition rather than a response to a problem.

Ideas, works, and debates

Works

Demonstrations rather than treatises are what moved the field.

  • Gender Shades (2018) measured commercial face-classification accuracy by skin tone and gender and found error rates an order of magnitude worse for darker-skinned women. It changed the field because it measured rather than argued.
  • “On the Dangers of Stochastic Parrots” (2021) sets out scale, environmental cost, training-data opacity and the coherence illusion — and the circumstances of its publication, with authors leaving the company involved, are part of what it demonstrated about institutional independence.
  • Model Cards and Datasheets for Datasets propose documentation as an accountability mechanism, which is the field's most practical output.
  • Weapons of Math Destruction (2016) and Race After Technology (2019) supply the public arguments; the fairness-metrics literature supplies the formal results, including the a guarantee that several intuitive fairness criteria cannot be satisfied simultaneously.

Central ideas

Four, and the fourth is what this work most has to answer to.

  • Measured disparity rather than asserted bias: the field's strongest work establishes differential performance empirically, which is what made it unanswerable in a way that argument had not been.
  • Fairness is not one thing, and the formal results show several reasonable definitions are mutually incompatible — so a system cannot be fair without someone choosing which fairness, and that choice is not technical.
  • Documentation as accountability: recording what a system was trained on, what it was evaluated for and where it fails is the mechanism that makes any later correction possible.
  • Fluency is not reliability. A system optimised to produce coherent output produces coherent output, and a reader's impression of understanding is evidence about the generation process rather than about what is being described.

Distinctive vocabulary

Three terms whose loose use is itself a problem the field studies.

  • Bias: in the technical literature a measurable disparity in performance or outcome across groups. Not prejudice, and using the word for both makes the measured claim sound like an accusation and the accusation sound measured.
  • Alignment: making a system's behaviour match a specified objective. Not making it good, and the gap between the specification and the intention is the actual problem.
  • Interpretability: the ability to say something about why a system produced an output. Not the system explaining itself, which is a different and less reliable thing.

Debates and disagreements

The field disagrees about its own subject, and about who is asking.

  • Present harms against long-range risk: whether the priority is measurable damage in deployed systems or the possibility of severe harm from more capable ones. The dispute is partly empirical and substantially about attention and funding, and it has become tribal in ways that serve neither side.
  • Whether principles work. Dozens of published AI-ethics principle sets converge on similar values and there is little evidence they change what gets built, which is the field's most uncomfortable finding about itself.
  • Institutional capture: much of the research is funded or employed by the organisations being evaluated, and the field has documented cases of that constraining what could be published.
  • Whether ethics is the right frame at all, or whether the questions are political and regulatory and calling them ethical is a way of keeping them internal.

Intellectual relationships

It inherits from fields that had these arguments earlier.

  • Philosophy of technology, documented separately here, supplies the finding that artifacts are not neutral — which AI ethics rediscovered rather than inherited, and often without the citation.
  • Cybernetics, also documented here, supplies the oldest version of the alignment problem: its founder wrote in the 1940s and 1960s that a machine given an objective pursues that objective and not the one intended.
  • Testimony theory and epistemic injustice, documented here, supply the vocabulary for what a system does when it allocates credibility.
  • Science and technology studies supplies the method for studying the organisations rather than only the models.
  • Fairness in machine learning is the formal wing and produces results the philosophical wing frequently does not read.

How it changes this reading

AI ethics carries thirteen placements here — ten claims and three through-lines — and it is not a reference this work can hold at arm's length.

  • This artifact is a machine-assisted reading of a person's life, published. The field's core finding about fluency applies to it directly: coherent, confident prose about someone is evidence about how the prose was generated and not about the someone.
  • Documentation as accountability is why the claims about version history, primary material and completed artifacts sit here. Recording what a reading was built from and where it fails is the mechanism that makes correction possible, and it is the field's most transferable practical result.
  • The incompatibility results supply a discipline the work needs: when reasonable criteria conflict, a system does not satisfy them all by being careful. Someone chooses, and the choice should be visible rather than dissolved into good intentions.
  • The uncomfortable finding is also inherited. Principle sets have not been shown to change what gets built, and a work that states its own ethical commitments at length is exactly the artifact that finding is about.
  • They are not evidence about how this or any system was built, evaluated, or used.

Useful comparisons

Four adjacent fields, three documented separately here.

  • AI ethics: measured disparity, documentation, and the governance of the organisations building the systems.
  • Philosophy of technology: artifacts distributing possibility, argued a generation earlier and at a different scale.
  • Cybernetics: the objective-given-versus-intended gap, stated first and largely unattributed since.
  • Media theory: what a medium's storage and retrieval make findable, which is the same question about a different layer.

Where the ideas meet

All four deny that a built system is a neutral instrument.

  • None treats a technology as indifferent to what passes through it.
  • All hold that design decisions distribute consequences beyond anyone's intention.
  • Each treats the conditions of production as part of the analysis rather than background to it.

Where they part

The newest field is the one least likely to cite the others.

  • Philosophy of technology and cybernetics reached several of AI ethics's central claims decades earlier, and the rediscovery is usually uncredited — which matters because the older fields also documented the failure modes of the responses.
  • AI ethics is entangled with the industry it studies in a way the older fields were not, so its literature has a funding structure that the others' analyses did not have to account for in themselves.
  • The formal fairness wing and the critical wing address different objects and rarely read each other, which is why the field can simultaneously hold that fairness is formally impossible and that fairness is a matter of power.

Limits

What these placements can carry here is narrow, and this reference has an unusual conflict.

  • The field's empirical results concern specific systems evaluated on specific tasks. They establish nothing about any other system, including this one.
  • Its principle literature has not been shown to change outcomes, so citing principles is not evidence that anything was done.
  • This work is inside the field's subject matter, which means its use of these references is not disinterested and cannot be presented as such.

Criticisms

The field's sharpest critics are inside it.

  • Principle proliferation without demonstrated effect is the standing internal criticism: convergent published values, little evidence of changed practice.
  • Institutional capture is documented rather than alleged, including cases where publication was constrained by the organisation employing the researchers.
  • The present-harms and long-range-risk camps have become tribal, and the tribalism has cost both sides evidence they would otherwise use.
  • Benchmarks and fairness metrics can be optimised against, which converts a measurement into a target and destroys it as a measurement.
  • The field frames as ethical a set of questions that are substantially political and regulatory, which keeps them inside the organisations rather than in front of the people affected.

Common misreadings

Four uses these placements do not license.

  • Using bias in its technical and its accusatory sense interchangeably, which borrows measurement for a charge and charge for a measurement.
  • Citing a stated principle as evidence that a system behaves accordingly, which is precisely what the field found does not follow.
  • Reading fluent output as evidence of understanding, which is the coherence illusion the field named.
  • Treating fairness as achievable by care, when the formal results show reasonable criteria conflict and someone has to choose.

What remains outside this idea

  • It cannot establish how any system was built, what it was trained on, or how it behaves.
  • It cannot convert a stated commitment into evidence that it was met.

References for further reading

Primary Source

“Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification,” Proceedings of Machine Learning Research 81 (2018); “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?” Proceedings of FAccT (2021); “Model Cards for Model Reporting,” Proceedings of FAT* (2019).

Secondary Source

“Inherent Trade-Offs in the Fair Determination of Risk Scores” (2017), which proves that several intuitive fairness criteria cannot be satisfied together; “The Global Landscape of AI Ethics Guidelines,” Nature Machine Intelligence 1 (2019), which surveys the principle sets and their convergence without demonstrated effect.