All posts

Why GPT-4 Failed Its Safety Test (and Passed It)

Dr. Jerry A. Smith · September 7, 2025 · 11 min read

Listen to the article on Apple Podcasts
Listen to the article on Soundcloud

Kant’s Phenomena, AI’s Noumena: Why Scientific Objectivity in AI Is a Cultural Construct

Last year, GPT-4 received opposite safety ratings from different evaluation teams using identical protocols. One declared it “sufficiently aligned,” another flagged “concerning capabilities.” Both followed rigorous procedures, yet reached contradictory conclusions. This wasn’t a failure of method — it revealed something fundamental about AI science: what we call “objectivity” is constructed through human categories, not discovered in nature.

Immanuel Kant provides the precise framework for understanding this. His insight: We never access reality directly; our minds structure only appearances. This doesn’t make science arbitrary — it makes it human. In AI, where we engineer every aspect from datasets to benchmarks, this constructed nature becomes impossible to ignore.

The Kantian Framework: Built-In Limits to Knowledge

Kant distinguished phenomena (things as they appear to us) from noumena (things as they exist independently). We can only know phenomena because our cognition actively structures experience through built-in categories, such as causality, substance, and quantity. These aren’t learned from experience — they’re what make experience possible (Kant, 1998/1781–1787).

Consider how we perceive causation. We don’t see “cause” directly; we see temporal sequences and impose causal structure. When a ball hits another ball, we experience succession, but our mind supplies the concept “A caused B to move.” This happens through schemata—procedures that bridge pure concepts and messy experiences.

For science, this means objectivity isn’t about escaping human perspective but about establishing intersubjective validity — shared rules that any rational observer can follow. Scientific knowledge is “objective” when it uses transparent, reproducible procedures, not when it achieves a view from nowhere (Kant, 1998).

Perception, Perspective, and the Active Mind

Kant revolutionized philosophy by arguing that perception is active, not passive. The mind doesn’t merely receive impressions — it constructs experience through synthesis. This occurs at multiple levels:

Perception involves the immediate structuring of sensory input. When we see a tree, we don’t receive raw green patches and brown streaks. Our perceptual apparatus automatically organizes these into a unified object persisting through time and space. Cognitive science confirms this Kantian insight: visual processing involves predictive coding, where the brain generates models and updates them based on sensory error signals (Clark, 2013).

Perspective adds another layer — the specific standpoint from which perception occurs. Two observers perceiving the same phenomenon may construct different experiences based on their conceptual frameworks, prior knowledge, and purposes. A botanist sees species indicators and ecological relationships; an artist sees form and color harmonies. Both perceive the same noumenal tree but construct different phenomenal realities.

This dual structuring — through universal perceptual categories and particular perspectival frameworks — explains why science achieves partial objectivity without neutrality. Researchers share basic perceptual and logical structures while bringing different theoretical perspectives, creating what Longino (2002) calls “pluralistic objectivity.”

Biology Shows How Science Embeds Culture

The history of reproductive biology illustrates how cultural categories shape “objective” observations. For decades, scientific literature described sperm as aggressive warriors racing to penetrate passive eggs — mirroring societal gender roles. The egg was portrayed as a damsel in distress, waiting for rescue. Martin (1991) documented how these metaphors pervaded textbooks and research articles, shaping hypotheses and interpretations.

Later research revealed the egg actively selects and signals to sperm, even digesting those deemed unsuitable. The shift from passive to active descriptions paralleled changing cultural views about gender. Critically, the earlier descriptions weren’t “wrong” — they captured certain phenomena while obscuring others. The perspective shaped what could be perceived (Fausto-Sterling, 2000).

Neuroscience exhibits similar patterns. The dominant “central command” model positions the brain as CEO and the body as an obedient workforce. But network models revealing bidirectional gut-brain communication and distributed decision-making challenge this hierarchy. The enteric nervous system contains 500 million neurons — more than the spinal cord — and operates quasi-independently. The model we choose affects what we study and what we discover (Fleck, 1979/1935).

These aren’t mere metaphors — they’re constitutive. They determine which questions get asked, what counts as evidence, and how anomalies get explained. As Daston and Galison (2007) documented, even the ideal of “objectivity” has evolved from “truth-to-nature” (selecting ideal specimens) to “mechanical objectivity” (letting instruments decide) to “trained judgment” (expert interpretation).

AI: Pure Phenomena Without Noumena

In AI, we never encounter a “mind-in-itself.” We only access behaviors, outputs, and measurements — phenomena structured by our engineering choices:

Representation (Kant’s “forms of intuition”): Tokenization determines what the model can perceive. GPT-4 literally cannot see that “running” and “running” are the same word if tokenized differently. A vision model at 224×224 pixels experiences a different visual world than one at 512×512. These aren’t neutral technical choices — they embed assumptions about what matters. Byte-pair encoding privileges common English patterns; character-level encoding might better capture morphology in agglutinative languages.

Architecture (Kant’s “categories”): Attention mechanisms, convolutions, and recurrence act as built-in categories determining what patterns can be learned. A transformer cannot discover specific sequential dependencies that RNNs naturally capture, just as humans cannot perceive ultraviolet without tools. The inductive biases we build into architectures shape what kinds of intelligence can emerge.

Training objectives (Kant’s “schemata”): Loss functions operationalize abstract concepts. When we train for “helpfulness,” we create temporal procedures that bridge the idea and model behavior. RLHF doesn’t discover helpfulness — it constructs it through specific rating procedures (Ouyang et al., 2022). Different annotation guidelines produce different models, each embodying a particular perspective on what helpfulness means.

Guiding ideals (Kant’s “regulative ideas”): Concepts like “AGI” or “alignment” aren’t empirical observations, but rather horizons that organize research programs. Teams pursuing AGI make different architectural choices than those building narrow tools. The perspective shapes the phenomenon.

Multiple Perspectives, Multiple AIs

The perspectival nature of AI becomes vivid when we examine how different stakeholders experience the same system:

Developers perceive AI through metrics — perplexity scores, benchmark results, and compute efficiency. Their phenomenon is statistical, optimizable, and controllable.

Users perceive AI through interactions — response quality, reliability, ease of use. Their phenomenon is pragmatic, experiential, and often surprising.

Affected communities perceive AI through impacts — job displacement, bias amplification, and surveillance potential. Their phenomenon is political, consequential, and often concerning.

Regulators perceive AI through risks — market concentration, safety failures, and misuse potential. Their phenomenon is systemic, requiring governance.

Each perspective reveals and conceals different aspects. A facial recognition system that developers perceive as “99% accurate” may be experienced by Black women as discriminatory (30% error rate), by privacy advocates as surveillance infrastructure, and by law enforcement as a crime-fighting tool. These aren’t disagreements about facts but different phenomenal constructions of the same noumenal system.

Where Values Enter AI’s “Objectivity”

Every stage of AI development embeds values disguised as technical choices:

Data curation: The Gender Shades study revealed commercial face classifiers failed disproportionately on darker-skinned women — a direct result of training data composition (Buolamwini & Gebru, 2018). “Representative” datasets encode decisions about which populations matter. ImageNet’s person categories included more fine-grained labels for dogs than for people from non-Western countries, shaping what models learned to distinguish (Crawford & Paglen, 2019).

Objective functions: Safety and capability aren’t natural categories with optimal trade-offs that can be discovered through experimentation. Tightening refusal rates reduces utility; loosening them increases risk. The “right” balance reflects social priorities, not physical constants. Constitutional AI (Bai et al., 2022) makes this explicit by encoding specific principles, but all AI systems embed implicit constitutions.

Evaluation: Benchmarks crystallize values. MMLU tests academic knowledge (Hendrycks et al., 2021), coding benchmarks test software engineering — both reasonable but partial views of “intelligence.” Models optimized for these metrics inherit their implicit definitions. The perspective embedded in the benchmark becomes the phenomenon measured.

Why This Matters: The Real Stakes of AI’s Constructed Objectivity

This isn’t merely an academic exercise. The myth of AI neutrality has concrete consequences that ripple through society:

Hidden values become embedded infrastructure. When we pretend AI is objective, unexamined choices calcify into systems that shape millions of lives. Consider hiring algorithms trained on “successful employee” data that encode historical discrimination. By treating the pattern as neutral rather than constructed, companies automate bias while claiming objectivity. Amazon’s scrapped recruiting tool that penalized resumes mentioning “women’s” colleges exemplifies this — the system wasn’t broken, it perfectly reflected its constructed definition of merit.

Power tends to concentrate around those who define “objectivity.” When AI development proceeds under the neutrality myth, whoever controls the benchmarks controls the field. Current definitions of AI progress — climbing leaderboards, scaling parameters — reflect specific visions of intelligence that favor certain architectures, applications, and actors. Alternative perspectives on intelligence (embodied, social, culturally embedded) struggle for legitimacy because they don’t register on established metrics.

Accountability becomes impossible. The neutrality myth creates a responsibility vacuum. When AI systems produce harmful outputs, developers invoke technical necessity: “We’re just optimizing the objective function.” Regulators often lack the technical vocabulary to contest specific choices effectively. Affected communities can’t challenge embedded values because they’re hidden behind claims of scientific objectivity. The construct of neutrality becomes a shield against democratic governance.

Real-world AI failures trace to perspective blindness. Predictive policing systems claim to identify crime risks objectively but actually map historical police deployment patterns, creating feedback loops that intensify surveillance of overpoliced communities. Medical AI trained predominantly on light-skinned patients misdiagnoses skin conditions in darker-skinned populations. Language models trained on internet text reproduce the perspectives of those with internet access and time to write. Each failure stems from mistaking a particular perspective for universal truth.

Robust Science Without the Neutrality Myth

This analysis might suggest relativism: if everything is constructed, is anything true? Kant’s answer was no — intersubjective validity remains possible through shared procedures. Modern science studies scholars like Longino and Haraway strengthen this: acknowledging standpoints improves objectivity by making hidden assumptions visible and testable.

Consider physics. Planes fly not because physics achieves a “view from nowhere” but because its procedures reliably identify invariances — patterns that hold across contexts. Similarly, robust AI science should pursue:

  • Transfer tests: Does performance generalize across distributions? (Recht et al., 2019; Taori et al., 2020)
  • Ablation studies: Which components causally contribute to capabilities?
  • Stakeholder audits: Do outputs align with affected communities’ values?
  • Perspective documentation: Whose viewpoint shaped this evaluation?

Haraway’s (1988) concept of “situated knowledges” offers a path forward. Rather than pursuing impossible neutrality, we should cultivate explicitly positioned, accountable perspectives. This isn’t relativism but what Barad (2007) calls “agential realism” — understanding that our measuring apparatuses (including conceptual frameworks) partially constitute what we measure.

A Method for AI’s Democratic Objectivity

If AI objectivity is constructed through choices, those choices should be transparent and contestable:

  1. Document value decisions: For each model, maintain public logs of dataset filters, objective weights, and safety thresholds with explicit justifications (Mitchell et al., 2019; Gebru et al., 2021).
  2. Map stakeholder perspectives: Identify whose phenomena matter — developers, users, affected communities — and how their perspectives differ.
  3. Diversify evaluation: Combine capability metrics with harm audits and distribution shift tests. Report not just average performance but variance across populations.
  4. State guiding ideals: Make explicit whether optimizing for “general intelligence,” “narrow assistance,” or “conservative safety.” Each implies different trade-offs.
  5. Enable contestation: Regular stakeholder reviews with published responses. Not consensus-seeking but assumption-revealing.
  6. Version value frameworks: As social priorities shift, document how and why evaluation criteria change.

This is “democratic” not because popularity determines truth, but because public procedures decide which claims are well-supported for shared action—a Kantian intersubjectivity updated for complex socio-technical systems.

Conclusion: From Neutrality Myth to Transparent Construction

Kant showed that our cognitive structures necessarily mediate human knowledge. AI makes this mediation visible and engineerable: we build the representations, objectives, and benchmarks that determine what our systems perceive and how we evaluate them.

Recognizing AI objectivity as constructed isn’t a weakness — it’s intellectually honest and practically powerful. Instead of pretending to achieve impossible neutrality, we can openly negotiate which values to embed, document our choices, and build systems whose biases are at least transparent and contestable.

When we stop claiming AI is objective in the mythical sense, we can start making it objective in Kant’s sense: through rigorous, transparent, intersubjective procedures that acknowledge their own conditions and limits. That’s not the death of scientific objectivity — it’s its maturation.

The path forward isn’t to abandon objectivity but to understand it correctly: as an achievement of transparent, contestable, shared procedures rather than a transcendent view from nowhere. Only then can AI development serve human flourishing rather than entrenching existing power structures under the guise of neutral science.

References

Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., … & Kaplan, J. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073.

Barad, K. (2007). Meeting the universe halfway: Quantum physics and the entanglement of matter and meaning. Duke University Press.

Buolamwini, J., & Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. Proceedings of Machine Learning Research, 81, 1–15.

Clark, A. (2013). Whatever next? Predictive brains, situated agents, and the future of cognitive science. Behavioral and Brain Sciences, 36(3), 181–204.

Crawford, K., & Paglen, T. (2019). Excavating AI: The politics of training sets for machine learning. AI Now Institute.

Daston, L., & Galison, P. (2007). Objectivity. Zone Books.

Fausto-Sterling, A. (2000). Sexing the body: Gender politics and the construction of sexuality. Basic Books.

Fleck, L. (1979). Genesis and development of a scientific fact (F. Bradley & T. J. Trenn, Trans.). University of Chicago Press. (Original work published 1935)

Gebru, T., Morgenstern, J., Vecchione, B., Vaughan, J. W., Wallach, H., Daumé III, H., & Crawford, K. (2021). Datasheets for datasets. Communications of the ACM, 64(12), 86–92.

Haraway, D. (1988). Situated knowledges: The science question in feminism and the privilege of partial perspective. Feminist Studies, 14(3), 575–599.

Hendrycks, D., Burns, C., Basart, S., Critch, A., Li, J., Song, D., & Steinhardt, J. (2021). Measuring massive multitask language understanding. Proceedings of the International Conference on Learning Representations (ICLR).

Kant, I. (1998). Critique of pure reason (P. Guyer & A. W. Wood, Trans.). Cambridge University Press. (Original work published 1781/1787)

Longino, H. E. (2002). The fate of knowledge. Princeton University Press.

Martin, E. (1991). The egg and the sperm: How science has constructed a romance based on stereotypical male-female roles. Signs, 16(3), 485–501.

Mitchell, M., Wu, S., Zaldivar, A., Barnes, P., Vasserman, L., Hutchinson, B., Spitzer, E., Raji, I. D., & Gebru, T. (2019). Model cards for model reporting. Proceedings of the Conference on Fairness, Accountability, and Transparency (FAccT), 220–229.

Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., … & Lowe, R. (2022). Training language models to follow instructions with human feedback. arXiv preprint arXiv:2203.02155.

Recht, B., Roelofs, R., Schmidt, L., & Shankar, V. (2019). Do ImageNet classifiers generalize to ImageNet? Advances in Neural Information Processing Systems (NeurIPS), 32, 1–12.

Taori, R., Dave, A., Shankar, V., Carlini, N., Wong, E., Coenen, A., … & Schmidt, L. (2020). Measuring robustness to natural distribution shifts in image classification. Advances in Neural Information Processing Systems (NeurIPS), 33, 18583–18599.

Start with one workflow.

Tell me what your team does today, where the work gets stuck, and what a useful result would look like. We will use a short call to identify a sensible next step.

Book a call