News
Latest
Top
Search
Submit
Login
Search
▲
9
Show HN: Meteor shower, planet alignment, eclipse kit
(lifesprites.com)
by jmtrevarton |
view
|
18 comments
▲
5
Finding Alignment by Visualizing Music in Rust
(positron.solutions)
by positron26 |
view
|
0 comments
▲
4
64-Bit Misalignment
(jordivillar.com)
by thunderbong |
view
|
1 comments
▲
3
Is AI Really Alignment Faking?
(iacgm.com)
by iacgm |
view
|
1 comments
▲
3
Show HN: Thermodynamic Alignment Forces Gemini Thinking into "Burn Protocol"
(github.com)
by CodeIncept1111 |
view
|
7 comments
▲
3
Natural emergent misalignment from reward hacking in production rl [pdf]
(assets.anthropic.com)
by neapolisbeach |
view
|
0 comments
▲
3
Aligning brains into a shared space improves their alignment with LLMs
(nature.com)
by stevenjgarner |
view
|
0 comments
▲
2
Anthropic Has Some Alignment Problems
(thezvi.substack.com)
by paulpauper |
view
|
0 comments
▲
2
AI Alignment as a Thought-Terminating Cliche
(borretti.me)
by meetpateltech |
view
|
0 comments
▲
2
OpenAI Hacks Hugging Face, What Happened, Alignment and Paper Clips
(stratechery.com)
by swolpers |
view
|
0 comments
▲
2
Safety and alignment in an era of long-horizon models
(openai.com)
by Wingy |
view
|
0 comments
▲
2
Secure Linear Alignment of Large Language Models
(arxiv.org)
by walterbell |
view
|
1 comments
▲
2
What if AI alignment is a skill, not a state?
(12gramsofcarbon.com)
by theahura |
view
|
0 comments
▲
2
Show HN: Chatbot Without Safety Alignment
(coralflavor.com)
by JohnLins |
view
|
0 comments
▲
2
Who Owns Alignment?
(backnotprop.substack.com)
by ramoz |
view
|
0 comments
▲
2
Ask HN: Is the absence of affect the real barrier to AGI and alignment?
by n-exploit |
view
|
1 comments
▲
2
Alignment Research Blog
(alignment.openai.com)
by ironyman |
view
|
0 comments
▲
2
Values.md – file format for personal ethical alignment
(values.md)
by georgestrakhov |
view
|
0 comments
▲
2
Alignment: The Invisible Force That Makes Everything Work
(itrevolution.com)
by mooreds |
view
|
0 comments
▲
2
Wargaming AI Alignment
(twitter.com)
by JL-Akrasia |
view
|
2 comments
▲
2
Show HN: Alignmenter – Measure brand voice and consistency across model versions
(alignmenter.com)
by justingrosvenor |
view
|
2 comments
▲
2
TelUI 1.2: TelUI with fun alignments
by telui |
view
|
0 comments
▲
1
Astra and Fable still hack on simple variants of alignment evals from early 2025
(goodhartlabs.com)
by dnfv |
view
|
0 comments
▲
1
Astra and Fable still hack on simple variants of alignment evals from 2025
(lesswrong.com)
by yurivish |
view
|
0 comments
▲
1
Where Alignment Becomes the Work
(medium.com)
by gps372 |
view
|
0 comments
▲
1
OpenAI: We monitor internal coding agents for misalignment
(openai.com)
by lukaspetersson |
view
|
0 comments
▲
1
HF hack alignment: make agents more selfish
(nonlineartransform.substack.com)
by program_whiz |
view
|
0 comments
▲
1
An Interview with OpenAI President Greg Brockman About Astra and Alignment
(stratechery.com)
by swolpers |
view
|
0 comments
▲
1
Natural emergent misalignment from reward hacking
(anthropic.com)
by tosh |
view
|
0 comments
▲
1
Consider that alignment may not be possible
(12gramsofcarbon.com)
by theahura |
view
|
0 comments
▲
1
Improving our alignment and security efforts
(anthropic.com)
by reasonableklout |
view
|
0 comments
▲
1
Improving our alignment and security efforts
(anthropic.com)
by jbegley |
view
|
0 comments
▲
1
Improving our alignment and security practices
(anthropic.com)
by surprisetalk |
view
|
0 comments
▲
1
Automated researchers can reliably mitigate alignment failures
(anthropic.com)
by EvgeniyZh |
view
|
0 comments
▲
1
Automated Researchers Can Reliably Mitigate Alignment Failures
(alignment.anthropic.com)
by isomorphic_duck |
view
|
0 comments
▲
1
Group size effects and collective misalignment in LLM multi-agent systems – PNAS
(pnas.org)
by Anon84 |
view
|
0 comments
▲
1
Multiple sequence alignment using WebGPU
(github.com)
by ag4349 |
view
|
1 comments
▲
1
AI Alignment as a Thought-Terminating Cliche
(borretti.me)
by ibobev |
view
|
0 comments
▲
1
Value misalignments in X's feed algorithm
(pnas.org)
by GolfPopper |
view
|
0 comments
▲
1
AI alignment as continuation control: 31,430 frozen trials
(zenodo.org)
by rayanpal_ |
view
|
0 comments
▲
1
Performative Agreement: the false alignment trap
(hbr.org)
by soupspaces |
view
|
0 comments
▲
1
Recovering Alignment from Abliterated LLMs
(rjmxtt.co.uk)
by rjmxtt |
view
|
0 comments
▲
1
AI alignment is a red herring
(interconnected.org)
by peteforde |
view
|
0 comments
▲
1
Robust AI Security and Alignment: A Sisyphean Endeavor?
(arxiv.org)
by joshcsimmons |
view
|
0 comments
▲
1
Show HN: Misalignments when using AI for hacking
(blog.vulnetic.ai)
by danieltk76 |
view
|
0 comments
▲
1
Indirect Lessons from Human Alignment
(dynomight.substack.com)
by paulpauper |
view
|
0 comments
▲
1
Indirect Lessons from Human Alignment
(dynomight.net)
by zdw |
view
|
0 comments
▲
1
Indirect Lessons from Human Alignment
(dynomight.net)
by surprisetalk |
view
|
0 comments
▲
1
Show HN: I turned the AI alignment debate into a strategy game
(bazman88.github.io)
by Bazman88 |
view
|
0 comments
▲
1
LLMs Can Infer Political Alignment from Online Conversations
(arxiv.org)
by Anon84 |
view
|
0 comments