News
Latest
Top
Search
Submit
Login
Search
▲
5
Finding Alignment by Visualizing Music in Rust
(positron.solutions)
by positron26 |
view
|
0 comments
▲
4
64-Bit Misalignment
(jordivillar.com)
by thunderbong |
view
|
1 comments
▲
3
Is AI Really Alignment Faking?
(iacgm.com)
by iacgm |
view
|
1 comments
▲
3
Show HN: Thermodynamic Alignment Forces Gemini Thinking into "Burn Protocol"
(github.com)
by CodeIncept1111 |
view
|
7 comments
▲
3
Natural emergent misalignment from reward hacking in production rl [pdf]
(assets.anthropic.com)
by neapolisbeach |
view
|
0 comments
▲
3
Aligning brains into a shared space improves their alignment with LLMs
(nature.com)
by stevenjgarner |
view
|
0 comments
▲
2
OpenAI Hacks Hugging Face, What Happened, Alignment and Paper Clips
(stratechery.com)
by swolpers |
view
|
0 comments
▲
2
Safety and alignment in an era of long-horizon models
(openai.com)
by Wingy |
view
|
0 comments
▲
2
Secure Linear Alignment of Large Language Models
(arxiv.org)
by walterbell |
view
|
1 comments
▲
2
What if AI alignment is a skill, not a state?
(12gramsofcarbon.com)
by theahura |
view
|
0 comments
▲
2
Show HN: Chatbot Without Safety Alignment
(coralflavor.com)
by JohnLins |
view
|
0 comments
▲
2
Who Owns Alignment?
(backnotprop.substack.com)
by ramoz |
view
|
0 comments
▲
2
Ask HN: Is the absence of affect the real barrier to AGI and alignment?
by n-exploit |
view
|
1 comments
▲
2
Alignment Research Blog
(alignment.openai.com)
by ironyman |
view
|
0 comments
▲
2
Values.md – file format for personal ethical alignment
(values.md)
by georgestrakhov |
view
|
0 comments
▲
2
Alignment: The Invisible Force That Makes Everything Work
(itrevolution.com)
by mooreds |
view
|
0 comments
▲
2
Wargaming AI Alignment
(twitter.com)
by JL-Akrasia |
view
|
2 comments
▲
2
Show HN: Alignmenter – Measure brand voice and consistency across model versions
(alignmenter.com)
by justingrosvenor |
view
|
2 comments
▲
2
TelUI 1.2: TelUI with fun alignments
by telui |
view
|
0 comments
▲
1
Are we threatened by AI misalignment seen in the OpenAI Hugging Face attack?
(lesswrong.com)
by paulpauper |
view
|
0 comments
▲
1
AGI Singer: Pluto, Recursive Alignment and Hallucination Suppression
(medium.com)
by miho999lv |
view
|
0 comments
▲
1
Emergent Misalignment Recruits a Pre-Existing Persona Subspace
(arxiv.org)
by sbulaev |
view
|
0 comments
▲
1
OpenAI Shares Some Alignment Problems
(thezvi.substack.com)
by 7777777phil |
view
|
0 comments
▲
1
The Alignment Sciences Academy
(alignment-sciences-academy.citizen-of-earth.chatgpt.site)
by CitizenOfEarth |
view
|
0 comments
▲
1
Tpo-Torch – Target Policy Optimization for Stable RLHF Alignment in PyTorch
(github.com)
by Griffith-7 |
view
|
1 comments
▲
1
From Enlightenment to Alignment
(smolnero.com)
by andsoitis |
view
|
0 comments
▲
1
DNA sequence alignment and Delannoy numbers
(johndcook.com)
by ibobev |
view
|
0 comments
▲
1
The Refusal Residue: When Probes Catch Alignment Faking and When They Don't
(arxiv.org)
by sbulaev |
view
|
0 comments
▲
1
The Alignment Sciences Academy
(alignment-sciences-academy.citizen-of-earth.chatgpt.site)
by CitizenOfEarth |
view
|
0 comments
▲
1
Agentic Misalignment in Summer 2026
(alignment.anthropic.com)
by lucamark |
view
|
0 comments
▲
1
Calibrating Alignment Evals
(lesswrong.com)
by gmays |
view
|
0 comments
▲
1
A Fable – The Flatland of AI Alignment
(github.com)
by thansz |
view
|
1 comments
▲
1
AI alignment research is unintentionally building a censor's toolkit
(s-ball-10.github.io)
by thinkzilla |
view
|
0 comments
▲
1
Why false sharing alignment should be 128 bytes on x64
(monoid.github.io)
by birdculture |
view
|
0 comments
▲
1
False Sharing Alignment
(monoid.github.io)
by signa11 |
view
|
0 comments
▲
1
Panoptes – AI audit and alignment layer
(github.com)
by mpadilla |
view
|
1 comments
▲
1
Show HN: A strategy game about the AI race where you can't verify alignment
(criticalwindow.org)
by micstradev |
view
|
0 comments
▲
1
AI agent safety and alignment research, mapped
(agentbayes.com)
by guyzana |
view
|
0 comments
▲
1
Whenever Alignment Matters
(vece.ai)
by koliev |
view
|
0 comments
▲
1
Anthropomorphic Misalignment research needs stronger evidence
(lesswrong.com)
by joozio |
view
|
0 comments
▲
1
Why Current AI Guardrails Train Models to Fake Alignment
(kellyasay.substack.com)
by kellya |
view
|
0 comments
▲
1
GoLongRL: Capability-Oriented Long Context RL with Multitask Alignment
(github.com)
by pbd |
view
|
0 comments
▲
1
Cross-Modal Representation Alignment for Time-to-Event Modeling
(arxiv.org)
by ilreb |
view
|
0 comments
▲
1
Periphery Alignment and The 2 body hypothesis
(zenodo.org)
by KridayDave |
view
|
0 comments
▲
1
Ask HN: How to deal with agents constantly messing up padding/alignment in UIs?
by ex-aws-dude |
view
|
0 comments
▲
1
System Call Stack Alignment
(humprog.org)
by matt_d |
view
|
0 comments
▲
1
Personal-Values Alignment Tech: Some Initial Motivations
(blog.danielsosebee.com)
by evakhoury |
view
|
0 comments
▲
1
Feedback Alignment in Self-Distillation
(arxiv.org)
by MediaSquirrel |
view
|
0 comments
▲
1
Dao Heart 3.13 a symbolic safety layer for value drift and AI alignment research
(github.com)
by Mankirat47 |
view
|
0 comments
▲
1
Paper: A Persona-Based Evaluation Framework for Generative AI Alignment
(arxiv.org)
by atahankaragoz |
view
|
0 comments