Dottor Anghel AI

Legal Counsel - Compliance Strategist in AI & Crypto | AI Act, GDPR, MiCA

I make AI adoptable, scalable, and defensible for organizations.

๐Ÿ’ช๐Ÿป My strength: I speak both the CTO's and General Counsel's language.

๐Ÿ” I help with:
โ€ข AI Governance & risk classification
โ€ข AI Act Audit & Readiness Assessment
โ€ข Generative AI Integration & internal automations
โ€ข Legal-tech support for IT, R&D, and Product teams
โ€ข M&A due diligence & algorithmic risk evaluation

๐Ÿ“Œ Why companies choose me:
โœ“ Operational output, not theory
โœ“ Metrics-driven, surgical approach
โœ“ AI compliance = competitive advantage
โœ“ Fast execution with precision

๐ŸŽฌ In this channel:
โ€ข How LLMs actually work
โ€ข AI Act compliance & regulatory deep dives
โ€ข Generative AI risks & mitigation strategies
โ€ข Critical perspectives on AI safety & algorithmic governance

๐ŸŽฏ Mission: Help companies use AI aggressively, safely, at scale โ€” eliminating regulatory risk and inefficiencies.

๐Ÿ‘ฅ For: General Counsels, CTOs, DPOs, Legal professionals


Dottor Anghel AI

โš ๏ธ ๐—–๐—ข๐—ก๐—ง๐—˜๐—ซ๐—ง ๐—ฅ๐—ข๐—ง ๐—œ๐—ฆ๐—ก'๐—ง ๐—” ๐—•๐—จ๐—š. ๐—œ๐—ง'๐—ฆ ๐—” ๐—™๐—จ๐—ก๐——๐—”๐— ๐—˜๐—ก๐—ง๐—”๐—Ÿ ๐—ฅ๐—œ๐—š๐—›๐—ง๐—ฆ ๐—ฅ๐—œ๐—ฆ๐—ž.

If you read my previous post on Ralph Loop, you know LLM performance degrades as context grows. Scientifically documented. Unavoidable.

But here's what nobody is talking about: context rot isn't a "performance issue." It's a systemic risk that should be in every FRIA under the AI Act.

And I've never seen it there.

๐—” ๐—ฅ๐—˜๐—”๐—Ÿ ๐—ฆ๐—–๐—˜๐—ก๐—”๐—ฅ๐—œ๐—ข

An HR platform uses AI to screen 500+ CVs daily.

9:00 AM - Candidate #1 (Mario):
โ†’ AI session starts fresh
โ†’ Context: 500 tokens
โ†’ Performance: optimal
โ†’ Score: 85/100 โ†’ hired

5:00 PM - Candidate #300 (Sara):
โ†’ Same AI session (8 hours running)
โ†’ Context: 45,000 tokens (299 CVs processed before her)
โ†’ Performance: degraded by context rot
โ†’ Score: 72/100 โ†’ rejected

Sara's CV was identical to Mario's. Same skills, same experience, same job posting.

She was rejected because the AI processed her last, not because she was less qualified.

That's not a bug. That's systematic discrimination based on processing order.

๐—ช๐—›๐—ฌ ๐—ง๐—›๐—œ๐—ฆ ๐—œ๐—ฆ ๐—” ๐—™๐—ฅ๐—œ๐—” ๐—ฃ๐—ฅ๐—ข๐—•๐—Ÿ๐—˜๐— 

FRIA requires identifying risks to non-discrimination and fair treatment.

"Continuous State" architectures (like Ralph Loop's second implementation) create this risk by design:
โŒ Performance varies by position in sequence
โŒ Early inputs get better treatment than late inputs
โŒ Order-based discrimination is architecturally guaranteed

"Clean State" architectures mitigate it:
โœ“ Context resets for each input
โœ“ Consistent performance across all candidates
โœ“ No order-based bias

One architecture creates fundamental rights risk. The other prevents it.

๐—ช๐—›๐—”๐—ง ๐—ง๐—›๐—œ๐—ฆ ๐— ๐—˜๐—”๐—ก๐—ฆ ๐—™๐—ข๐—ฅ ๐—–๐—ข๐— ๐—ฃ๐—Ÿ๐—œ๐—”๐—ก๐—–๐—˜

If you're conducting FRIA for AI systems:

๐—ค๐—จ๐—˜๐—ฆ๐—ง๐—œ๐—ข๐—ก ๐Ÿญ: Does your system accumulate context over time?
If yes โ†’ you have documented performance degradation risk.

๐—ค๐—จ๐—˜๐—ฆ๐—ง๐—œ๐—ข๐—ก ๐Ÿฎ: Does processing order affect outcomes?
If "Continuous State" โ†’ the answer is mathematically yes.

๐—ค๐—จ๐—˜๐—ฆ๐—ง๐—œ๐—ข๐—ก ๐Ÿฏ: Have you documented context rot as a fundamental rights risk?
If no โ†’ your FRIA is incomplete.

๐—ง๐—›๐—˜ ๐—Ÿ๐—œ๐—”๐—•๐—œ๐—Ÿ๐—œ๐—ง๐—ฌ ๐—ค๐—จ๐—˜๐—ฆ๐—ง๐—œ๐—ข๐—ก

When Sara files a GDPR complaint showing identical CVs but different outcomes, who's liable?

โ†’ The vendor, for choosing vulnerable architecture?
โ†’ The deployer, for not identifying this systemic risk?
โ†’ Both?

The AI Act doesn't answer this yet. But auditors will ask.

๐—›๐—ข๐—ช ๐—œ ๐—”๐—ฃ๐—ฃ๐—ฅ๐—ข๐—”๐—–๐—› ๐—ง๐—›๐—œ๐—ฆ

I don't conduct "generic" FRIAs based on use case alone.

I assess architectural risks:
โœ“ Does the architecture create systematic performance variance?
โœ“ Can processing order introduce discrimination?
โœ“ Is context rot documented and mitigated?
โœ“ Are vendor contracts clear on liability for architecture-induced bias?

Because an HR AI that discriminates based on when you applied (not what you submitted) violates fundamental rights.

Even if nobody intended it. Even if "it just happened" because of context rot.

Intent doesn't matter. Impact does.

๐Ÿ“ฉ Conducting FRIA for AI systems with autonomous agents? dott.anghel.ai@gmail.com

#AIAct #FRIA #FundamentalRights #ContextRot #AIGovernance #Compliance #Discrimination #TechLaw #RalphLoop #AIAgents

7 months ago | [YT] | 3

Dottor Anghel AI

๐ŸŽญ ๐—ฅ๐—”๐—Ÿ๐—ฃ๐—› ๐—ช๐—œ๐—š๐—š๐—จ๐—  ๐—”๐—ก๐—— ๐—ง๐—›๐—˜ ๐—”๐—œ ๐—”๐—š๐—˜๐—ก๐—ง ๐—ง๐—›๐—”๐—ง ๐—ช๐—œ๐—ก๐—ฆ ๐—•๐—ฌ ๐—ก๐—˜๐—ฉ๐—˜๐—ฅ ๐—ค๐—จ๐—œ๐—ง๐—ง๐—œ๐—ก๐—š

In 2024, Geoffrey Huntley built an AI coding agent named after the least competent character from The Simpsons: Ralph Wiggum, the kid who said "I'm learnding!" and ate paste.

The joke was perfect. Ralph doesn't win by being smart. It wins by being persistent.

๐—ง๐—›๐—˜ ๐—ข๐—ฅ๐—œ๐—š๐—œ๐—ก๐—”๐—Ÿ: ๐—ข๐—ก๐—˜ ๐—Ÿ๐—œ๐—ก๐—˜ ๐—ข๐—™ ๐—•๐—”๐—ฆ๐—›

while :; do cat PROMPT.md | claude-code ; done

Loop forever. Feed the prompt to Claude. If it fails, try again.

The philosophy: eventual consistency. Try enough times, the AI will produce working code.

It built complete projects, a new programming language, and six production repos in one night during a Y Combinator hackathon.

๐—ง๐—›๐—˜ ๐—˜๐—ฉ๐—ข๐—Ÿ๐—จ๐—ง๐—œ๐—ข๐—ก: ๐—ง๐—ช๐—ข ๐—ฃ๐—›๐—œ๐—Ÿ๐—ข๐—ฆ๐—ข๐—ฃ๐—›๐—œ๐—˜๐—ฆ

As Ralph gained traction, two camps emerged:

๐—–๐—Ÿ๐—˜๐—”๐—ก ๐—ฆ๐—ง๐—”๐—ง๐—˜ (snarktank/ralph):
Kill the session after every task. Start fresh.
โ†’ New Claude instance each iteration
โ†’ Context always minimal
โ†’ State persists externally: git, progress.txt, prd.json

๐—–๐—ข๐—ก๐—ง๐—œ๐—ก๐—จ๐—ข๐—จ๐—ฆ ๐—ฆ๐—ง๐—”๐—ง๐—˜ (Claude Code plugin):
Keep the session alive. Loop within one conversation.
โ†’ One long session, never terminates
โ†’ Context accumulates indefinitely
โ†’ State persists in model's memory
โ†’ Requires circuit breakers, timeouts, limits

Same goal. Opposite execution.

๐—ง๐—›๐—˜ ๐—ฃ๐—ฅ๐—ข๐—•๐—Ÿ๐—˜๐— : ๐—–๐—ข๐—ก๐—ง๐—˜๐—ซ๐—ง ๐—ฅ๐—ข๐—ง

LLMs don't process token 10,000 like token 100. Performance degrades as context grows. Always.

Chroma's research documented "context rot":
โ†’ Longer context = worse performance
โ†’ Past information becomes distractors
โ†’ Hallucination rates increase
โ†’ Models get less reliable over time

LongMemEval proved it: focused input outperforms full context, even when full context contains more information.

๐—ช๐—›๐—ฌ ๐—œ๐—ง ๐— ๐—”๐—ง๐—ง๐—˜๐—ฅ๐—ฆ

"Clean State" is architecturally aligned with context rot. It fights the problem by design.

"Continuous State" is architecturally vulnerable. It needs complex scaffolding to prevent performance collapse.

One works with the science. The other against it.

๐—ง๐—›๐—˜ ๐—š๐—ข๐—ฉ๐—˜๐—ฅ๐—ก๐—”๐—ก๐—–๐—˜ ๐—ค๐—จ๐—˜๐—ฆ๐—ง๐—œ๐—ข๐—ก

If you're deploying autonomous AI agents:

โ†’ Can you document which architecture your system uses?
โ†’ Have you assessed context rot as a risk factor?
โ†’ Can you prove reliability doesn't degrade over time?

When a regulator asks "how does your AI maintain consistent performance?", "it just works" isn't documentation.

Understanding whether your system works with or against fundamental LLM limitations is.

Ralph Wiggum succeeded because it embraced persistence over perfection. Your AI governance should do the same: build systems that work with the technology's limits, not against them.

๐Ÿ“ฉ Building or procuring autonomous AI systems? dott.anghel.ai@gmail.com

#AI #AIGovernance #RalphLoop #LLM #ContextRot #AIAgents #Compliance #MachineLearning #AIAct #TechLaw

7 months ago | [YT] | 3

Dottor Anghel AI

๐ŸŽจ ๐—›๐—ข๐—ช ๐—”๐—œ "๐—œ๐— ๐—”๐—š๐—œ๐—ก๐—˜๐—ฆ" ๐—ช๐—›๐—”๐—ง ๐——๐—ข๐—˜๐—ฆ๐—ก'๐—ง ๐—˜๐—ซ๐—œ๐—ฆ๐—ง: ๐—ง๐—ช๐—ข ๐— ๐—ข๐——๐—˜๐—Ÿ๐—ฆ ๐—–๐—ข๐— ๐—ฃ๐—”๐—ฅ๐—˜๐——

When you ask Midjourney or DALL-E to generate an image, what really happens under the hood?

The answer isn't just one. There are two completely different philosophies for "creating from nothing":

๐Ÿญ. ๐——๐—œ๐—™๐—™๐—จ๐—ฆ๐—œ๐—ข๐—ก ๐— ๐—ข๐——๐—˜๐—Ÿ๐—ฆ (the "clearing noise" method)
How it works: start from pure random noise and gradually "clean it up" until you get the desired image
โœ… Pros: exceptional photographic quality, fine detail control
โŒ Cons: computationally expensive, requires many iterative steps
๐Ÿ“Œ Examples: Stable Diffusion, DALL-E, Midjourney

๐Ÿฎ. ๐—”๐—จ๐—ง๐—ข๐—ฅ๐—˜๐—š๐—ฅ๐—˜๐—ฆ๐—ฆ๐—œ๐—ฉ๐—˜ ๐— ๐—ข๐——๐—˜๐—Ÿ๐—ฆ (the "pixel by pixel" method)
How it works: generates the image one token at a time, just like ChatGPT writes text word by word
โœ… Pros: unified text+image architecture, scalability, speed in certain contexts
โŒ Cons: historically inferior quality for realistic images
๐Ÿ“Œ Examples: DALL-E 3 (hybrid), visual transformer models

๐—•๐—จ๐—ง ๐—ง๐—›๐—˜๐—ฅ๐—˜'๐—ฆ ๐—” ๐—ก๐—˜๐—ช ๐—œ๐——๐—˜๐—” ๐—ง๐—›๐—”๐—ง ๐—–๐—›๐—”๐—ก๐—š๐—˜๐—ฆ ๐—˜๐—ฉ๐—˜๐—ฅ๐—ฌ๐—ง๐—›๐—œ๐—ก๐—š

Researchers are developing hybrid models like ๐—š๐—Ÿ๐— -๐—œ๐—บ๐—ฎ๐—ด๐—ฒ that combine the best of both worlds:
โ†’ The efficiency of autoregressive
โ†’ The quality of diffusion
โ†’ A single architecture for text, image, video

In the video, I show you with clear animations:
โœ“ How diffusion "removes noise" step by step
โœ“ How autoregressive "builds" the image token by token
โœ“ Why GLM-Image could be the future of multimodal generation

๐—ช๐—›๐—ฌ ๐—ฆ๐—›๐—ข๐—จ๐—Ÿ๐—— ๐—ฌ๐—ข๐—จ ๐—ž๐—ก๐—ข๐—ช ๐—ง๐—›๐—˜๐—ฆ๐—˜ ๐——๐—œ๐—™๐—™๐—˜๐—ฅ๐—˜๐—ก๐—–๐—˜๐—ฆ?

If you work with generative AI (or are considering implementing it), knowing which model your tool uses changes:
โ€ข The performance you can expect
โ€ข The actual computational costs
โ€ข The technical (and legal) limits of generation
โ€ข How to document the process for AI Act compliance

Understanding the architecture isn't just technical curiosity: it's conscious governance.

๐Ÿ“ฝ๏ธ Watch the full video to see the differences in action (6 min) -> https://www.youtube.com/watch?v=B-pYz...

#AI #GenerativeAI #ImageGeneration #AIAct #Diffusion #Transformer #TechEducation #MachineLearning #AIGovernance #DeepLearning

7 months ago | [YT] | 6

Dottor Anghel AI

โ™€๏ธ ๐—”๐—œ ๐—”๐—ก๐—— ๐—ฉ๐—œ๐—ข๐—Ÿ๐—˜๐—ก๐—–๐—˜ ๐—”๐—š๐—”๐—œ๐—ก๐—ฆ๐—ง ๐—ช๐—ข๐— ๐—˜๐—ก: ๐— ๐—ฌ ๐—œ๐—ก๐—ง๐—˜๐—ฅ๐—ฉ๐—˜๐—ก๐—ง๐—œ๐—ข๐—ก ๐—”๐—ง ๐—ง๐—›๐—˜ ๐—–๐—›๐—”๐— ๐—•๐—˜๐—ฅ ๐—ข๐—™ ๐——๐—˜๐—ฃ๐—จ๐—ง๐—œ๐—˜๐—ฆ

๐Ÿ“ Yesterday, at the Chamber of Deputies, I had the honour of addressing one of the most urgent and troubling challenges of our time.

The new frontier of violence against women increasingly takes the form of digital abuse, identity manipulation and reputational harm enabled by artificial intelligence and deepfake technologies.

As a jurist and expert in artificial intelligence and emerging technologies, I delivered an intervention entitled โ€œAI, Deepfakes and the New Frontier of Violence Against Womenโ€, focusing on:

๐Ÿ‡ช๐Ÿ‡บ The European regulatory framework (AI Act, Directive (EU) 2024/1385, Digital Services Act)

๐Ÿ‡ฎ๐Ÿ‡น The Italian legal response, with Law no. 132/2025 and the introduction of Article 612-quater of the Italian Criminal Code, criminalising the dissemination of deepfake content.

๐—œ๐—ก๐—ฆ๐—ง๐—œ๐—ง๐—จ๐—ง๐—œ๐—ข๐—ก๐—”๐—Ÿ ๐—”๐—–๐—ž๐—ก๐—ข๐—ช๐—Ÿ๐—˜๐——๐—š๐—˜๐— ๐—˜๐—ก๐—ง๐—ฆ

I wish to express my sincere gratitude to:

๐Ÿ‡ท๐Ÿ‡ดย H.E. Gabriela Dancau, Ambassador of Romania to Italy, San Marino and Malta,

๐Ÿ‡ธ๐Ÿ‡ฒ H.E. Marina Emiliani, Ambassador of the Republic of San Marino to Romania,

for their presence and sensitivity towards a phenomenon that deeply affects womenโ€™s dignity, honour and identity.

My thanks also go to Hon. Luciano Ciocchetti, organiser of the initiative, and to Hon. Martina Semenzato, President of the Parliamentary Commission of Inquiry on Femicide and all forms of gender-based violence, for making this institutional discussion possible.

๐—•๐—˜๐—ฌ๐—ข๐—ก๐—— ๐—Ÿ๐—˜๐—š๐—œ๐—ฆ๐—Ÿ๐—”๐—ง๐—œ๐—ข๐—ก

I would also like to thank Gianluca Mech, creator of the short film โ€œLa Trappola di Venereโ€, for his contribution to a cultural and preventive intervention addressed to men, aimed at recognising psychological and emotional warning signs before distress turns into violence.

The myth of Venus serves as a metaphor for the loss of clarity caused by obsessive love.

Events like this show that effective protection requires action on multiple levels: legal, technological, educational and cultural.

#AI #Deepfakes #ViolenceAgainstWomen #AIAct #DigitalLaw #AIGovernance #GenderBasedViolence

7 months ago | [YT] | 7

Dottor Anghel AI

๐Ÿ“Œ ๐—ฅ๐—˜๐—ฃ๐—˜๐—ง๐—œ๐—ง๐—” ๐—œ๐—จ๐—ฉ๐—”๐—ก๐—ง: ๐—ช๐—›๐—ฌ ๐——๐—จ๐—ฃ๐—Ÿ๐—œ๐—–๐—”๐—ง๐—œ๐—ก๐—š ๐—ฌ๐—ข๐—จ๐—ฅ ๐—ฃ๐—ฅ๐—ข๐— ๐—ฃ๐—ง ๐— ๐—œ๐—š๐—›๐—ง ๐—•๐—˜ ๐—ง๐—›๐—˜ ๐—ฆ๐—œ๐— ๐—ฃ๐—Ÿ๐—˜๐—ฆ๐—ง ๐—ช๐—”๐—ฌ ๐—ง๐—ข ๐—•๐—ข๐—ข๐—ฆ๐—ง ๐—Ÿ๐—Ÿ๐—  ๐—ฃ๐—˜๐—ฅ๐—™๐—ข๐—ฅ๐— ๐—”๐—ก๐—–๐—˜

Google Research just published a study confirming what your grandmother always told you: repetition works.

And it's almost embarrassingly simple.

Take your prompt. Copy-paste it twice. That's it. No complex Chain-of-Thought. No advanced engineering. Just: Prompt + Prompt = Better Results.

๐—ช๐—›๐—ฌ ๐——๐—ข๐—˜๐—ฆ ๐—ง๐—›๐—œ๐—ฆ ๐—ช๐—ข๐—ฅ๐—ž?

LLMs are causal models: they process tokens left-to-right and can't "look ahead".

โœ… Single prompt โ†’ the model processes each token as it arrives, with limited context.

โœ… Duplicated prompt โ†’ when the model reaches the second repetition, it can attend to the full prompt from the first pass.

It's like saying: "Read this carefully, form a complete picture, and NOW answer me."

๐—ง๐—›๐—˜ ๐—ฅ๐—˜๐—ฆ๐—จ๐—Ÿ๐—ง๐—ฆ ๐—”๐—ฅ๐—˜ ๐—ฆ๐—ง๐—”๐—š๐—š๐—˜๐—ฅ๐—œ๐—ก๐—š

Google tested this across 70 benchmarks (ARC, GSM8K, MMLU-Pro, MATH, etc.):

โœ…
- 47 wins out of 70 tests
- 0 losses (yes, zero)
- Accuracy jumps from 21.33% to 97.33% in some tasks
- Works across Gemini, GPT-4o, Claude 3.5, DeepSeek
- No latency increase: pre-fill phase is parallel, so response time stays the same

โŒ
- Only applies to non-reasoning models (standard inference, not o1-style deliberation)
- Effect varies by task and model architecture

๐—ช๐—›๐—”๐—ง ๐—ง๐—›๐—œ๐—ฆ ๐— ๐—˜๐—”๐—ก๐—ฆ ๐—™๐—ข๐—ฅ ๐—ฃ๐—ฅ๐—ข๐——๐—จ๐—–๐—ง๐—œ๐—ข๐—ก ๐—ฆ๐—ฌ๐—ฆ๐—ง๐—˜๐— ๐—ฆ

For teams running production LLMs, this is a zero-cost performance upgrade:

1. Drop-in implementation: modify your prompt wrapper to duplicate input
2. No infrastructure changes: same API, same latency
3. Immediate gains on structured tasks: reasoning, Q&A, classification

But here's the governance angle most teams miss:

If you're documenting AI system behavior for compliance (AI Act Art. 13, transparency requirements), prompt repetition gives you a reproducible, explainable intervention you can defend to auditors.

You're not changing the model. You're not injecting opaque vectors. You're justโ€ฆ giving it more time to think. Auditors love simplicity.

๐—ช๐—›๐—˜๐—ก ๐—ง๐—ข ๐—จ๐—ฆ๐—˜ ๐—œ๐—ง (๐—”๐—ก๐—— ๐—ช๐—›๐—˜๐—ก ๐—ก๐—ข๐—ง ๐—ง๐—ข)

โœ… USE IT FOR:
- Structured reasoning tasks (math, logic, multi-step questions)
- High-stakes outputs where accuracy > cost
- Non-reasoning models (GPT-4, Claude 3.5, Gemini 2.0)

โŒ DON'T USE IT FOR:
- Reasoning models (o1, DeepSeek-R1) โ†’ they already do internal iteration
- Creative/open-ended generation โ†’ repetition doesn't help much
- Latency-critical applications where token count matters

Sometimes we overcomplicate things. We chase RAG pipelines, fine-tuning, complex prompt chains.

And the solution is literally: say it twice.

๐Ÿ“„ Link to the full paper in the first comment

๐Ÿ“ฉ If you want to integrate simple, auditable prompt strategies into your AI compliance stack, contact me at: dott.anghel.ai@gmail.com

#AI #PromptEngineering #LLM #AIAct #Compliance #AIGovernance #GoogleResearch #FRIA

7 months ago | [YT] | 7

Dottor Anghel AI

๐Ÿ“Œ ๐—ก๐—˜๐—จ๐—ฅ๐—ข๐—ก๐—ฃ๐—˜๐——๐—œ๐—”: ๐—ง๐—›๐—˜ โ€œ๐—ช๐—œ๐—ž๐—œ๐—ฃ๐—˜๐——๐—œ๐—” ๐—ข๐—™ ๐—ก๐—˜๐—จ๐—ฅ๐—ข๐—ก๐—ฆโ€ ๐—™๐—ข๐—ฅ ๐—Ÿ๐—Ÿ๐— ๐—ฆ

In modern LLMs we talk about โ€œneuronsโ€ and โ€œfeaturesโ€, but most teams never see what they actually do.
Neuronpedia tries to fix that by turning neuron/SAE interpretability into a shared, navigable resource.

1๏ธโƒฃ ๐—ช๐—›๐—”๐—ง ๐—ก๐—˜๐—จ๐—ฅ๐—ข๐—ก๐—ฃ๐—˜๐——๐—œ๐—” ๐—œ๐—ฆ

A public, collaborative atlas of neurons and SAE features for real models.

PRO โœ…
- Explanations, example prompts, and activations for individual neurons/features.
- Strong focus on sparse autoencoders (SAEs) and their interpretable โ€œfeaturesโ€.
- Web UI + APIs so you can browse, tag, and analyse features without building your own tooling.

2๏ธโƒฃ ๐—ช๐—›๐—”๐—ง ๐—ฌ๐—ข๐—จ ๐—–๐—”๐—ก ๐——๐—ข ๐—ช๐—œ๐—ง๐—› ๐—œ๐—ง

Think of Neuronpedia as your starting point for mechโ€‘interp and steering.

PRO โœ…
- Inspect concrete features like โ€œlegal toneโ€, โ€œEiffel Towerโ€, or โ€œselfโ€‘harmโ€ with real examples.
- Compare features across layers/models and plug SAEs into existing analysis libraries.

CON โŒ (๐—ถ๐—ณ ๐˜†๐—ผ๐˜‚ ๐—ถ๐—ด๐—ป๐—ผ๐—ฟ๐—ฒ ๐—ถ๐˜)
- You keep treating steering vectors as blackโ€‘box magic instead of grounded, documented concepts.

3๏ธโƒฃ ๐—ฆ๐—ง๐—˜๐—˜๐—ฅ๐—œ๐—ก๐—š & ๐—ฆ๐—”๐—™๐—˜๐—ง๐—ฌ ๐—”๐—ก๐—š๐—Ÿ๐—˜

Neuronpedia is also a bridge between interpretability and realโ€‘world control.

- Use SAE features as steering knobs: upโ€‘ or downโ€‘weight a feature and see how outputs change.
- Design safety/style steering grounded in labeled features, not opaque directions, which reduces sideโ€‘effects and helps with governance.

If you care about serious governance, this is a big upgrade over โ€œwe changed the system prompt and it looks betterโ€.

4๏ธโƒฃ ๐—ช๐—›๐—ฌ ๐—ฃ๐—ฅ๐—”๐—–๐—ง๐—œ๐—ง๐—œ๐—ข๐—ก๐—˜๐—ฅ๐—ฆ ๐—ฆ๐—›๐—ข๐—จ๐—Ÿ๐—— ๐—–๐—”๐—ฅ๐—˜

- Cuts timeโ€‘toโ€‘experiment: hosting, visualisation, and collaboration are handled for you.
- Makes interventions easier to justify to stakeholders and regulators: you can point to specific, named features with examples, not just โ€œprompt engineering that seems to workโ€.

๐Ÿ“ฉ If you want to explore how Neuronpedia and steering vectors can fit into your AI governance or compliance stack, contact me at: dott.anghel.ai@gmail.com

#AI #Compliance #FRIA #AIAct #Neuronpedia

8 months ago | [YT] | 7

Dottor Anghel AI

๐Ÿ“Œ ๐— ๐—ข๐—ฅ๐—˜ ๐—™๐—˜๐—”๐—ง๐—จ๐—ฅ๐—˜๐—ฆ, ๐— ๐—ข๐—ฅ๐—˜ ๐—–๐—ข๐—ก๐—ง๐—ฅ๐—ข๐—Ÿ? ๐—ก๐—ข๐—ง ๐—ฅ๐—˜๐—”๐—Ÿ๐—Ÿ๐—ฌ. ๐—ฆ๐—œ๐—ก๐—š๐—Ÿ๐—˜โ€‘๐—™๐—˜๐—”๐—ง๐—จ๐—ฅ๐—˜ ๐—ฆ๐—ง๐—˜๐—˜๐—ฅ๐—œ๐—ก๐—š ๐—ข๐—™๐—ง๐—˜๐—ก ๐—ช๐—œ๐—ก๐—ฆ

In the SAE/steering world it sounds intuitive: โ€œmore features = more controlโ€.

Recent work suggests something else: more poorly chosen features can actually reduce coherence and control.

1๏ธโƒฃ ๐—ฆ๐—œ๐—ก๐—š๐—Ÿ๐—˜โ€‘๐—™๐—˜๐—”๐—ง๐—จ๐—ฅ๐—˜ ๐—ฆ๐—ง๐—˜๐—˜๐—ฅ๐—œ๐—ก๐—š: ๐—ฆ๐—ก๐—œ๐—ฃ๐—˜๐—ฅ ๐— ๐—ข๐——๐—˜

With an SAE, a single good feature is often enough to steer a behavior: refusal, legal tone, marketing style, etc.

โœ…

1. Clear, interpretable effect

2. Less interference with other behaviors

3. Easier to test, document, and defend in audits

โŒ (when you mix too many features)

1. Some mostly โ€œreadโ€ the input instead of driving the output

2. Vectors add up and introduce noise โ†’ text becomes less coherent, control less predictable

Moral: one wellโ€‘chosen feature > many โ€œvibesโ€‘basedโ€ features.

2๏ธโƒฃ ๐— ๐—จ๐—Ÿ๐—ง๐—œโ€‘๐—Ÿ๐—˜๐—ฉ๐—˜๐—Ÿ ๐—ฆ๐—ง๐—˜๐—˜๐—ฅ๐—œ๐—ก๐—š: ๐—ฃ๐—ข๐—ง๐—˜๐—ก๐—ง๐—œ๐—”๐—Ÿ, ๐—•๐—จ๐—ง ๐—ก๐—ข๐—ง ๐— ๐—”๐—š๐—œ๐—–

Intervening on multiple layers / intermediate levels only helps if you know what you are doing.

โœ…

1. Some layers are better for content (middle), others for style/output (late)

2. Picking 1โ€“3 โ€œkeyโ€ layers often beats โ€œsteer everywhereโ€

โŒ

There is no strong evidence that โ€œapplying the same feature on many layersโ€ systematically beats using it where it has the most impact:

you risk more complexity, more noise, and governance explanations nobody really believes.

3๏ธโƒฃ ๐—ง๐—›๐—˜ ๐—ฅ๐—˜๐—”๐—Ÿ ๐—ง๐—ฅ๐—œ๐—–๐—ž: ๐—ฅ๐—œ๐—š๐—›๐—ง ๐—ฉ๐—˜๐—–๐—ง๐—ข๐—ฅ + ๐—ช๐—˜๐—Ÿ๐—Ÿโ€‘๐—ง๐—จ๐—ก๐—˜๐—— ๐——๐—˜๐—–๐—ข๐——๐—œ๐—ก๐—š

In practical tests (including Eiffelโ€‘Towerโ€‘style demos), the quality jump does not come from โ€œmore featuresโ€ but from:

1. Clamping

Limit how much the steering vector can distort activations โ†’ the concept stays, with less overshooting and fewer weird outputs.

2. Lower temperature

Reduce sampling temperature โ†’ less randomness, better instruction following, the steering signal is not washed out.

3. Light repetition penalty

Apply a moderate repetition penalty โ†’ the concept remains present without obsessive repetition.

Net result: compared to โ€œsimple addition + standard samplingโ€, the combo single wellโ€‘chosen feature + clamping + lower T + light repetition penalty often yields outputs that are more coherent, better aligned, and still strongly express the target concept.


โ˜… BONUS โ†’ The same steering coefficient is not universal: its effect can swing a lot with different prompts and contexts. Recent work shows that for a fixed coefficient some inputs get strong steering, others weak or even inverted effects, so you must treat it as a promptโ€‘ and taskโ€‘dependent hyperparameter.



---

If you care about serious governance, the question is not โ€œhow many features can I turn on?โ€, but:

โ€œWhich few features can I actually explain, control, and document โ€“ and how do I set up decoding around them?โ€

#AIAct #Compliance #Steering #Feature #AI #LegalTech

8 months ago | [YT] | 7

Dottor Anghel AI

๐Ÿ“Œ ๐—ฃ๐—ฅ๐—ข๐— ๐—ฃ๐—ง ๐—˜๐—ก๐—š๐—œ๐—ก๐—˜๐—˜๐—ฅ๐—œ๐—ก๐—š ๐—ฉ๐—ฆ ๐—ฆ๐—ง๐—˜๐—˜๐—ฅ๐—œ๐—ก๐—š: ๐—ช๐—›๐—ข ๐—ฆ๐—›๐—ข๐—จ๐—Ÿ๐—— ๐—–โ€‘๐—Ÿ๐—˜๐—ฉ๐—˜๐—Ÿ๐—ฆ ๐—•๐—˜๐—ง ๐—ข๐—ก?

Many companies are hiring prompt engineers.
Very few are building steering infrastructure.

If you need to govern models in production, this isnโ€™t a โ€œstyleโ€ choice โ€“ itโ€™s a choice about control, risk, and scalability.

1๏ธโƒฃ PROMPT ENGINEERING

โœ… PROS
1. No model access needed: the API is enough.
2. Fast iteration: you can tweak prompts in minutes.
3. Nonโ€‘technical teams can contribute.

โŒ CONS
1. Fragile: small changes in context can break behavior.
2. Hard to scale: each team โ€œinventsโ€ its own prompts, so crossโ€‘product / crossโ€‘country consistency is weak.
3. Auditing is painful: itโ€™s hard to explain to an auditor why a specific output appeared.

2๏ธโƒฃ STEERING

โœ… PROS
1. Internal control: you act on hidden states with concept vectors (risk, tone, safety, etc.).
2. Continuous adjustment: X_next = X + ฮฑ V โ†’ you dial behavior up or down instead of rewriting prompts.
3. Central governance: you can define โ€œcorporateโ€ vector libraries and reuse them across products and use cases.

โŒ CONS
1. Needs lowโ€‘level access (selfโ€‘hosted models or advanced tooling).
2. Technical setup: extracting, testing, and documenting robust vectors is not a side project.
3. Not every behavior is a โ€œsingle directionโ€: some vectors are less reliable and must be refined.


3๏ธโƒฃ HOW TO IMPROVE STEERING: FROM โ€œJUST ADDโ€ TO FINEโ€‘GRAIN CONTROL

The basic scheme is:
- addition: X_next = X + ฮฑ V
- standard sampling: default temperature, no extra tuning.

Practical tests show you can make steering much more stable with:

1. Clamping
Limit how much the vector can deform activations (per dimension, norm, or token).
Effect: the concept stays, but you avoid overshooting and weird outputs.

2. Lower temperature
Reduce sampling temperature.
Effect: less randomness, stronger adherence to instructions, the steering vector isnโ€™t washed out by noise.

3. Light repetition penalty
Apply a moderate repetition penalty.
Effect: the target concept remains present without obsessive, repetitive outputs.

Net result: compared to โ€œaddition + standard samplingโ€, the combo clamping + lower T + light repetition penalty produces smoother responses, better instructionโ€‘following, and equal or better preservation of the target concept.


For prototypes and lowโ€‘risk use cases, prompt engineering is enough.
For critical, auditable, multiโ€‘team products, you need to move from โ€œwriting promptsโ€ to designing real steering levers.

#AIAct #AI #PromptEngineering #Steering #Compliance

8 months ago | [YT] | 7

Dottor Anghel AI

๐Ÿงฉ ๐—ฆ๐—ฃ๐—”๐—ฅ๐—ฆ๐—˜ ๐—”๐—จ๐—ง๐—ข๐—˜๐—ก๐—–๐—ข๐——๐—˜๐—ฅ๐—ฆ (๐—ฆ๐—”๐—˜): ๐—ง๐—›๐—˜ ๐—›๐—œ๐——๐——๐—˜๐—ก ๐— ๐—”๐—ฃ ๐—ฌ๐—ข๐—จ ๐—ฆ๐—›๐—ข๐—จ๐—Ÿ๐—— ๐—•๐—˜ ๐—จ๐—ฆ๐—œ๐—ก๐—š ๐—ง๐—ข ๐—š๐—ข๐—ฉ๐—˜๐—ฅ๐—ก ๐—ฌ๐—ข๐—จ๐—ฅ ๐—Ÿ๐—Ÿ๐— ๐—ฆ

Most companies control AI with prompts and policies.
SAEs let you control it with internal switches, instead of hoping the model โ€œbehavesโ€.

Steering tells you how to push a model.
Concept vectors tell you in which direction.
Sparse Autoencoders (SAE) tell you what is actually inside, in a readable way.


1๏ธโƒฃ ๐—ช๐—›๐—”๐—ง ๐—”๐—ก ๐—ฆ๐—”๐—˜ ๐——๐—ข๐—˜๐—ฆ (๐—ฆ๐—œ๐— ๐—ฃ๐—Ÿ๐—œ๐—™๐—œ๐—˜๐——)

An SAE takes an LLMโ€™s activations (hidden states) and recodes them into:
1. a larger LATENT space
2. that is SPARSE: almost all values are zero.

Forcing sparsity has a key effect:
each active feature tends to represent a cleaner, more interpretable concept (or a small cluster of nearby concepts).

Instead of โ€œpolysemanticโ€ neurons doing a bit of everything, you get a list of switches:
1. โ€œlegaleseโ€
2. โ€œinformal languageโ€
3. โ€œhate speechโ€
4. โ€œcode / snippetsโ€
etc.

In practice, the SAE turns the chaos of activations into a switchboard, where each button turns on a specific behavior of the model.

And because there are thousands of features, you can use LLMs themselves to help name each switch (โ€œthis feature fires on phrases X: looks like โ€˜polite toneโ€™, โ€˜legaleseโ€™, etc.โ€) โ€“ this is autoโ€‘interpretability.

2๏ธโƒฃ ๐—›๐—ข๐—ช ๐—ฆ๐—”๐—˜๐—ฆ, ๐—–๐—ข๐—ก๐—–๐—˜๐—ฃ๐—ง ๐—ฉ๐—˜๐—–๐—ง๐—ข๐—ฅ๐—ฆ & ๐—ฆ๐—ง๐—˜๐—˜๐—ฅ๐—œ๐—ก๐—š ๐—™๐—œ๐—ง ๐—ง๐—ข๐—š๐—˜๐—ง๐—›๐—˜๐—ฅ

Each sparse SAE feature has:
1. a vector in latent space (the โ€œshapeโ€ of the concept)
2. an activation weight (how strongly it is on for that token / sequence).

This means you can:
1. Identify which features fire on certain concepts (toxicity, PII leaks, nonโ€‘compliant tone).
2. Treat that feature as a concept vector: the direction in latent space representing that behavior.

From there you go back to the steering formula:

X_next = X + V * ฮฑ

- X = current internal state of the model
- V = vector of the SAE feature (the concept direction)
- ฮฑ = how much you amplify or damp that concept

Itโ€™s the same mechanism as steering vectors: you press an internal button and push the modelโ€™s behavior in that direction.

3๏ธโƒฃ ๐—ช๐—›๐—ฌ ๐—Ÿ๐—˜๐—š๐—”๐—Ÿ / ๐—–๐—ง๐—ข / ๐—ฅ๐—œ๐—ฆ๐—ž ๐—ฆ๐—›๐—ข๐—จ๐—Ÿ๐—— ๐—–๐—”๐—ฅ๐—˜

For AI governance, SAEs matter because they let you:
1. Map โ€œrisk zonesโ€ inside the model (features that light up on problematic content).
2. Build stable technical controls: not just prompts and policies, but mathematical levers over internal concepts.

Itโ€™s the shift from:
โ€œWe told the model not to do itโ€
to
โ€œWe directly lowered the activation of the features that cause that behaviorโ€.

โžก๏ธ ๐—ก๐—˜๐—ซ๐—ง ๐—ฃ๐—ข๐—ฆ๐—ง: ๐—ฃ๐—ฅ๐—ข๐— ๐—ฃ๐—ง ๐—˜๐—ก๐—š๐—œ๐—ก๐—˜๐—˜๐—ฅ๐—œ๐—ก๐—š ๐—ฉ๐—ฆ ๐—ฆ๐—ง๐—˜๐—˜๐—ฅ๐—œ๐—ก๐—š
Which one actually gives you control at scale? Weโ€™ll compare costs, risks, and tradeโ€‘offs from a Cโ€‘level and controlโ€‘function perspective, not a โ€œprompt hackerโ€ one.

#AIAct #AI #SAE

8 months ago (edited) | [YT] | 7

Dottor Anghel AI

๐Ÿง  ๐—–๐—ข๐—ก๐—–๐—˜๐—ฃ๐—ง ๐—ฉ๐—˜๐—–๐—ง๐—ข๐—ฅ๐—ฆ: ๐—ช๐—›๐—˜๐—ก ๐—” ๐—–๐—ข๐—ก๐—–๐—˜๐—ฃ๐—ง ๐—•๐—˜๐—–๐—ข๐— ๐—˜๐—ฆ ๐—” ๐— ๐—”๐—ง๐—›๐—˜๐— ๐—”๐—ง๐—œ๐—–๐—”๐—Ÿ ๐——๐—œ๐—ฅ๐—˜๐—–๐—ง๐—œ๐—ข๐—ก

In previous posts we talked about:

- steering = changing an LLMโ€™s behavior โ€œon the flyโ€
- concept embeddings & activation space = the internal map you steer on

Now comes the key piece: CONCEPT VECTORS.

A concept vector is a DIRECTION in activation space that corresponds to a specific concept:

โ€œlegal toneโ€, โ€œmarketing styleโ€, โ€œtoxicityโ€, โ€œrisk conservatismโ€.

๐Ÿ“Œ ๐—™๐—ฟ๐—ผ๐—บ ๐—ต๐—ถ๐—ฑ๐—ฑ๐—ฒ๐—ป ๐˜€๐˜๐—ฎ๐˜๐—ฒ ๐˜๐—ผ โ€œ๐˜€๐˜๐—ฒ๐—ฒ๐—ฟ๐—ฒ๐—ฑโ€ ๐˜€๐˜๐—ฎ๐˜๐—ฒ

Imagine a hidden state vector of an LLM at some layer:

Hidden State Vector:

X = [1; 3; 6; 7]

Now suppose weโ€™ve found a concept vector:

V = [โˆ’1; 0; 2; 1] (example)

When we do steering, we build a new state:

X_next = X + V * ๐œถ

where ฮฑ is a coefficient that controls how strongly we apply that concept (the โ€œintensity knobโ€).

In practice:

- X = how the model was โ€œthinkingโ€ before
- V = the direction โ€œmore legalโ€, โ€œless toxicโ€, โ€œmore conservativeโ€
- ๐œถ = the dial: 0.2, 0.5, 1.5โ€ฆ how hard you push the concept

Mathematically, itโ€™s just vector addition.

Operationally, itโ€™s a personality/behavior shift without touching the modelโ€™s weights.

โš™๏ธ ๐—›๐—ข๐—ช ๐—ช๐—˜ ๐—™๐—œ๐—ก๐—— ๐—ง๐—›๐—˜ ๐—ฉ๐—˜๐—–๐—ง๐—ข๐—ฅ ๐—ฉ

Two main (simplified) paths:

1๏ธโƒฃ Difference between two prompt groups

- Group A: prompts that EXPRESS the concept (e.g. toxic answers).
- Group B: prompts that DO NOT express it (e.g. neutral answers).

You compute the mean activations for A and for B, then:

V โ‰ˆ mean_activations(A) โˆ’ mean_activations(B)

This V points in the direction โ€œmore like A, less like Bโ€.

2๏ธโƒฃ Sparse Autoencoders (SAE) โ€“ teaser

Sparse Autoencoders take the modelโ€™s activations and rewrite them into a larger but SPARSE space (mostly zeros).

Each sparse โ€œfeatureโ€ often corresponds to an internal concept (or nearโ€‘concept): an interpretable pattern of behavior.

In short:

- the SAE gives you a โ€œbuttonโ€ (a feature) that turns a concept on/off
- that button corresponds to a vector in latent space you can use to steer the model

SAEs deserve their own post, so weโ€™ll go deeper next time.

๐Ÿงญ ๐—ช๐—›๐—ฌ ๐—–๐—ข๐—ก๐—–๐—˜๐—ฃ๐—ง ๐—ฉ๐—˜๐—–๐—ง๐—ข๐—ฅ๐—ฆ ๐— ๐—”๐—ง๐—ง๐—˜๐—ฅ ๐—™๐—ข๐—ฅ ๐—š๐—ข๐—ฉ๐—˜๐—ฅ๐—ก๐—”๐—ก๐—–๐—˜

They let you turn:

- โ€œthis model is too aggressive/creativeโ€

into

- โ€œI apply +0.3 on prudence, โˆ’0.2 on creativity in its internal activationsโ€.

Itโ€™s the shift from written policies to mathematical levers: controllable, measurable, reproducible.

โžก๏ธ ๐—ก๐—˜๐—ซ๐—ง ๐—ฃ๐—ข๐—ฆ๐—ง: Sparse Autoencoders and how to use them to discover and control hidden concepts in your LLMs (without manual reverseโ€‘engineering).

๐Ÿ“ฉ Want help mapping concept vectors for risk, tone, or safety in your own AI systems โ€“ and turning them into real governance levers, not just prompts? Write to me confidentially: dott.anghel.ai@gmail.com

#AI #AIAct #FRIA #ConceptVectors

8 months ago (edited) | [YT] | 7