Legal Counsel - Compliance Strategist in AI & Crypto | AI Act, GDPR, MiCA
I make AI adoptable, scalable, and defensible for organizations.
๐ช๐ป My strength: I speak both the CTO's and General Counsel's language.
๐ I help with:
โข AI Governance & risk classification
โข AI Act Audit & Readiness Assessment
โข Generative AI Integration & internal automations
โข Legal-tech support for IT, R&D, and Product teams
โข M&A due diligence & algorithmic risk evaluation
๐ Why companies choose me:
โ Operational output, not theory
โ Metrics-driven, surgical approach
โ AI compliance = competitive advantage
โ Fast execution with precision
๐ฌ In this channel:
โข How LLMs actually work
โข AI Act compliance & regulatory deep dives
โข Generative AI risks & mitigation strategies
โข Critical perspectives on AI safety & algorithmic governance
๐ฏ Mission: Help companies use AI aggressively, safely, at scale โ eliminating regulatory risk and inefficiencies.
๐ฅ For: General Counsels, CTOs, DPOs, Legal professionals
Dottor Anghel AI
โ ๏ธ ๐๐ข๐ก๐ง๐๐ซ๐ง ๐ฅ๐ข๐ง ๐๐ฆ๐ก'๐ง ๐ ๐๐จ๐. ๐๐ง'๐ฆ ๐ ๐๐จ๐ก๐๐๐ ๐๐ก๐ง๐๐ ๐ฅ๐๐๐๐ง๐ฆ ๐ฅ๐๐ฆ๐.
If you read my previous post on Ralph Loop, you know LLM performance degrades as context grows. Scientifically documented. Unavoidable.
But here's what nobody is talking about: context rot isn't a "performance issue." It's a systemic risk that should be in every FRIA under the AI Act.
And I've never seen it there.
๐ ๐ฅ๐๐๐ ๐ฆ๐๐๐ก๐๐ฅ๐๐ข
An HR platform uses AI to screen 500+ CVs daily.
9:00 AM - Candidate #1 (Mario):
โ AI session starts fresh
โ Context: 500 tokens
โ Performance: optimal
โ Score: 85/100 โ hired
5:00 PM - Candidate #300 (Sara):
โ Same AI session (8 hours running)
โ Context: 45,000 tokens (299 CVs processed before her)
โ Performance: degraded by context rot
โ Score: 72/100 โ rejected
Sara's CV was identical to Mario's. Same skills, same experience, same job posting.
She was rejected because the AI processed her last, not because she was less qualified.
That's not a bug. That's systematic discrimination based on processing order.
๐ช๐๐ฌ ๐ง๐๐๐ฆ ๐๐ฆ ๐ ๐๐ฅ๐๐ ๐ฃ๐ฅ๐ข๐๐๐๐
FRIA requires identifying risks to non-discrimination and fair treatment.
"Continuous State" architectures (like Ralph Loop's second implementation) create this risk by design:
โ Performance varies by position in sequence
โ Early inputs get better treatment than late inputs
โ Order-based discrimination is architecturally guaranteed
"Clean State" architectures mitigate it:
โ Context resets for each input
โ Consistent performance across all candidates
โ No order-based bias
One architecture creates fundamental rights risk. The other prevents it.
๐ช๐๐๐ง ๐ง๐๐๐ฆ ๐ ๐๐๐ก๐ฆ ๐๐ข๐ฅ ๐๐ข๐ ๐ฃ๐๐๐๐ก๐๐
If you're conducting FRIA for AI systems:
๐ค๐จ๐๐ฆ๐ง๐๐ข๐ก ๐ญ: Does your system accumulate context over time?
If yes โ you have documented performance degradation risk.
๐ค๐จ๐๐ฆ๐ง๐๐ข๐ก ๐ฎ: Does processing order affect outcomes?
If "Continuous State" โ the answer is mathematically yes.
๐ค๐จ๐๐ฆ๐ง๐๐ข๐ก ๐ฏ: Have you documented context rot as a fundamental rights risk?
If no โ your FRIA is incomplete.
๐ง๐๐ ๐๐๐๐๐๐๐๐ง๐ฌ ๐ค๐จ๐๐ฆ๐ง๐๐ข๐ก
When Sara files a GDPR complaint showing identical CVs but different outcomes, who's liable?
โ The vendor, for choosing vulnerable architecture?
โ The deployer, for not identifying this systemic risk?
โ Both?
The AI Act doesn't answer this yet. But auditors will ask.
๐๐ข๐ช ๐ ๐๐ฃ๐ฃ๐ฅ๐ข๐๐๐ ๐ง๐๐๐ฆ
I don't conduct "generic" FRIAs based on use case alone.
I assess architectural risks:
โ Does the architecture create systematic performance variance?
โ Can processing order introduce discrimination?
โ Is context rot documented and mitigated?
โ Are vendor contracts clear on liability for architecture-induced bias?
Because an HR AI that discriminates based on when you applied (not what you submitted) violates fundamental rights.
Even if nobody intended it. Even if "it just happened" because of context rot.
Intent doesn't matter. Impact does.
๐ฉ Conducting FRIA for AI systems with autonomous agents? dott.anghel.ai@gmail.com
#AIAct #FRIA #FundamentalRights #ContextRot #AIGovernance #Compliance #Discrimination #TechLaw #RalphLoop #AIAgents
7 months ago | [YT] | 3
View 0 replies
Dottor Anghel AI
๐ญ ๐ฅ๐๐๐ฃ๐ ๐ช๐๐๐๐จ๐ ๐๐ก๐ ๐ง๐๐ ๐๐ ๐๐๐๐ก๐ง ๐ง๐๐๐ง ๐ช๐๐ก๐ฆ ๐๐ฌ ๐ก๐๐ฉ๐๐ฅ ๐ค๐จ๐๐ง๐ง๐๐ก๐
In 2024, Geoffrey Huntley built an AI coding agent named after the least competent character from The Simpsons: Ralph Wiggum, the kid who said "I'm learnding!" and ate paste.
The joke was perfect. Ralph doesn't win by being smart. It wins by being persistent.
๐ง๐๐ ๐ข๐ฅ๐๐๐๐ก๐๐: ๐ข๐ก๐ ๐๐๐ก๐ ๐ข๐ ๐๐๐ฆ๐
while :; do cat PROMPT.md | claude-code ; done
Loop forever. Feed the prompt to Claude. If it fails, try again.
The philosophy: eventual consistency. Try enough times, the AI will produce working code.
It built complete projects, a new programming language, and six production repos in one night during a Y Combinator hackathon.
๐ง๐๐ ๐๐ฉ๐ข๐๐จ๐ง๐๐ข๐ก: ๐ง๐ช๐ข ๐ฃ๐๐๐๐ข๐ฆ๐ข๐ฃ๐๐๐๐ฆ
As Ralph gained traction, two camps emerged:
๐๐๐๐๐ก ๐ฆ๐ง๐๐ง๐ (snarktank/ralph):
Kill the session after every task. Start fresh.
โ New Claude instance each iteration
โ Context always minimal
โ State persists externally: git, progress.txt, prd.json
๐๐ข๐ก๐ง๐๐ก๐จ๐ข๐จ๐ฆ ๐ฆ๐ง๐๐ง๐ (Claude Code plugin):
Keep the session alive. Loop within one conversation.
โ One long session, never terminates
โ Context accumulates indefinitely
โ State persists in model's memory
โ Requires circuit breakers, timeouts, limits
Same goal. Opposite execution.
๐ง๐๐ ๐ฃ๐ฅ๐ข๐๐๐๐ : ๐๐ข๐ก๐ง๐๐ซ๐ง ๐ฅ๐ข๐ง
LLMs don't process token 10,000 like token 100. Performance degrades as context grows. Always.
Chroma's research documented "context rot":
โ Longer context = worse performance
โ Past information becomes distractors
โ Hallucination rates increase
โ Models get less reliable over time
LongMemEval proved it: focused input outperforms full context, even when full context contains more information.
๐ช๐๐ฌ ๐๐ง ๐ ๐๐ง๐ง๐๐ฅ๐ฆ
"Clean State" is architecturally aligned with context rot. It fights the problem by design.
"Continuous State" is architecturally vulnerable. It needs complex scaffolding to prevent performance collapse.
One works with the science. The other against it.
๐ง๐๐ ๐๐ข๐ฉ๐๐ฅ๐ก๐๐ก๐๐ ๐ค๐จ๐๐ฆ๐ง๐๐ข๐ก
If you're deploying autonomous AI agents:
โ Can you document which architecture your system uses?
โ Have you assessed context rot as a risk factor?
โ Can you prove reliability doesn't degrade over time?
When a regulator asks "how does your AI maintain consistent performance?", "it just works" isn't documentation.
Understanding whether your system works with or against fundamental LLM limitations is.
Ralph Wiggum succeeded because it embraced persistence over perfection. Your AI governance should do the same: build systems that work with the technology's limits, not against them.
๐ฉ Building or procuring autonomous AI systems? dott.anghel.ai@gmail.com
#AI #AIGovernance #RalphLoop #LLM #ContextRot #AIAgents #Compliance #MachineLearning #AIAct #TechLaw
7 months ago | [YT] | 3
View 1 reply
Dottor Anghel AI
๐จ ๐๐ข๐ช ๐๐ "๐๐ ๐๐๐๐ก๐๐ฆ" ๐ช๐๐๐ง ๐๐ข๐๐ฆ๐ก'๐ง ๐๐ซ๐๐ฆ๐ง: ๐ง๐ช๐ข ๐ ๐ข๐๐๐๐ฆ ๐๐ข๐ ๐ฃ๐๐ฅ๐๐
When you ask Midjourney or DALL-E to generate an image, what really happens under the hood?
The answer isn't just one. There are two completely different philosophies for "creating from nothing":
๐ญ. ๐๐๐๐๐จ๐ฆ๐๐ข๐ก ๐ ๐ข๐๐๐๐ฆ (the "clearing noise" method)
How it works: start from pure random noise and gradually "clean it up" until you get the desired image
โ Pros: exceptional photographic quality, fine detail control
โ Cons: computationally expensive, requires many iterative steps
๐ Examples: Stable Diffusion, DALL-E, Midjourney
๐ฎ. ๐๐จ๐ง๐ข๐ฅ๐๐๐ฅ๐๐ฆ๐ฆ๐๐ฉ๐ ๐ ๐ข๐๐๐๐ฆ (the "pixel by pixel" method)
How it works: generates the image one token at a time, just like ChatGPT writes text word by word
โ Pros: unified text+image architecture, scalability, speed in certain contexts
โ Cons: historically inferior quality for realistic images
๐ Examples: DALL-E 3 (hybrid), visual transformer models
๐๐จ๐ง ๐ง๐๐๐ฅ๐'๐ฆ ๐ ๐ก๐๐ช ๐๐๐๐ ๐ง๐๐๐ง ๐๐๐๐ก๐๐๐ฆ ๐๐ฉ๐๐ฅ๐ฌ๐ง๐๐๐ก๐
Researchers are developing hybrid models like ๐๐๐ -๐๐บ๐ฎ๐ด๐ฒ that combine the best of both worlds:
โ The efficiency of autoregressive
โ The quality of diffusion
โ A single architecture for text, image, video
In the video, I show you with clear animations:
โ How diffusion "removes noise" step by step
โ How autoregressive "builds" the image token by token
โ Why GLM-Image could be the future of multimodal generation
๐ช๐๐ฌ ๐ฆ๐๐ข๐จ๐๐ ๐ฌ๐ข๐จ ๐๐ก๐ข๐ช ๐ง๐๐๐ฆ๐ ๐๐๐๐๐๐ฅ๐๐ก๐๐๐ฆ?
If you work with generative AI (or are considering implementing it), knowing which model your tool uses changes:
โข The performance you can expect
โข The actual computational costs
โข The technical (and legal) limits of generation
โข How to document the process for AI Act compliance
Understanding the architecture isn't just technical curiosity: it's conscious governance.
๐ฝ๏ธ Watch the full video to see the differences in action (6 min) -> https://www.youtube.com/watch?v=B-pYz...
#AI #GenerativeAI #ImageGeneration #AIAct #Diffusion #Transformer #TechEducation #MachineLearning #AIGovernance #DeepLearning
7 months ago | [YT] | 6
View 1 reply
Dottor Anghel AI
โ๏ธ ๐๐ ๐๐ก๐ ๐ฉ๐๐ข๐๐๐ก๐๐ ๐๐๐๐๐ก๐ฆ๐ง ๐ช๐ข๐ ๐๐ก: ๐ ๐ฌ ๐๐ก๐ง๐๐ฅ๐ฉ๐๐ก๐ง๐๐ข๐ก ๐๐ง ๐ง๐๐ ๐๐๐๐ ๐๐๐ฅ ๐ข๐ ๐๐๐ฃ๐จ๐ง๐๐๐ฆ
๐ Yesterday, at the Chamber of Deputies, I had the honour of addressing one of the most urgent and troubling challenges of our time.
The new frontier of violence against women increasingly takes the form of digital abuse, identity manipulation and reputational harm enabled by artificial intelligence and deepfake technologies.
As a jurist and expert in artificial intelligence and emerging technologies, I delivered an intervention entitled โAI, Deepfakes and the New Frontier of Violence Against Womenโ, focusing on:
๐ช๐บ The European regulatory framework (AI Act, Directive (EU) 2024/1385, Digital Services Act)
๐ฎ๐น The Italian legal response, with Law no. 132/2025 and the introduction of Article 612-quater of the Italian Criminal Code, criminalising the dissemination of deepfake content.
๐๐ก๐ฆ๐ง๐๐ง๐จ๐ง๐๐ข๐ก๐๐ ๐๐๐๐ก๐ข๐ช๐๐๐๐๐๐ ๐๐ก๐ง๐ฆ
I wish to express my sincere gratitude to:
๐ท๐ดย H.E. Gabriela Dancau, Ambassador of Romania to Italy, San Marino and Malta,
๐ธ๐ฒ H.E. Marina Emiliani, Ambassador of the Republic of San Marino to Romania,
for their presence and sensitivity towards a phenomenon that deeply affects womenโs dignity, honour and identity.
My thanks also go to Hon. Luciano Ciocchetti, organiser of the initiative, and to Hon. Martina Semenzato, President of the Parliamentary Commission of Inquiry on Femicide and all forms of gender-based violence, for making this institutional discussion possible.
๐๐๐ฌ๐ข๐ก๐ ๐๐๐๐๐ฆ๐๐๐ง๐๐ข๐ก
I would also like to thank Gianluca Mech, creator of the short film โLa Trappola di Venereโ, for his contribution to a cultural and preventive intervention addressed to men, aimed at recognising psychological and emotional warning signs before distress turns into violence.
The myth of Venus serves as a metaphor for the loss of clarity caused by obsessive love.
Events like this show that effective protection requires action on multiple levels: legal, technological, educational and cultural.
#AI #Deepfakes #ViolenceAgainstWomen #AIAct #DigitalLaw #AIGovernance #GenderBasedViolence
7 months ago | [YT] | 7
View 0 replies
Dottor Anghel AI
๐ ๐ฅ๐๐ฃ๐๐ง๐๐ง๐ ๐๐จ๐ฉ๐๐ก๐ง: ๐ช๐๐ฌ ๐๐จ๐ฃ๐๐๐๐๐ง๐๐ก๐ ๐ฌ๐ข๐จ๐ฅ ๐ฃ๐ฅ๐ข๐ ๐ฃ๐ง ๐ ๐๐๐๐ง ๐๐ ๐ง๐๐ ๐ฆ๐๐ ๐ฃ๐๐๐ฆ๐ง ๐ช๐๐ฌ ๐ง๐ข ๐๐ข๐ข๐ฆ๐ง ๐๐๐ ๐ฃ๐๐ฅ๐๐ข๐ฅ๐ ๐๐ก๐๐
Google Research just published a study confirming what your grandmother always told you: repetition works.
And it's almost embarrassingly simple.
Take your prompt. Copy-paste it twice. That's it. No complex Chain-of-Thought. No advanced engineering. Just: Prompt + Prompt = Better Results.
๐ช๐๐ฌ ๐๐ข๐๐ฆ ๐ง๐๐๐ฆ ๐ช๐ข๐ฅ๐?
LLMs are causal models: they process tokens left-to-right and can't "look ahead".
โ Single prompt โ the model processes each token as it arrives, with limited context.
โ Duplicated prompt โ when the model reaches the second repetition, it can attend to the full prompt from the first pass.
It's like saying: "Read this carefully, form a complete picture, and NOW answer me."
๐ง๐๐ ๐ฅ๐๐ฆ๐จ๐๐ง๐ฆ ๐๐ฅ๐ ๐ฆ๐ง๐๐๐๐๐ฅ๐๐ก๐
Google tested this across 70 benchmarks (ARC, GSM8K, MMLU-Pro, MATH, etc.):
โ
- 47 wins out of 70 tests
- 0 losses (yes, zero)
- Accuracy jumps from 21.33% to 97.33% in some tasks
- Works across Gemini, GPT-4o, Claude 3.5, DeepSeek
- No latency increase: pre-fill phase is parallel, so response time stays the same
โ
- Only applies to non-reasoning models (standard inference, not o1-style deliberation)
- Effect varies by task and model architecture
๐ช๐๐๐ง ๐ง๐๐๐ฆ ๐ ๐๐๐ก๐ฆ ๐๐ข๐ฅ ๐ฃ๐ฅ๐ข๐๐จ๐๐ง๐๐ข๐ก ๐ฆ๐ฌ๐ฆ๐ง๐๐ ๐ฆ
For teams running production LLMs, this is a zero-cost performance upgrade:
1. Drop-in implementation: modify your prompt wrapper to duplicate input
2. No infrastructure changes: same API, same latency
3. Immediate gains on structured tasks: reasoning, Q&A, classification
But here's the governance angle most teams miss:
If you're documenting AI system behavior for compliance (AI Act Art. 13, transparency requirements), prompt repetition gives you a reproducible, explainable intervention you can defend to auditors.
You're not changing the model. You're not injecting opaque vectors. You're justโฆ giving it more time to think. Auditors love simplicity.
๐ช๐๐๐ก ๐ง๐ข ๐จ๐ฆ๐ ๐๐ง (๐๐ก๐ ๐ช๐๐๐ก ๐ก๐ข๐ง ๐ง๐ข)
โ USE IT FOR:
- Structured reasoning tasks (math, logic, multi-step questions)
- High-stakes outputs where accuracy > cost
- Non-reasoning models (GPT-4, Claude 3.5, Gemini 2.0)
โ DON'T USE IT FOR:
- Reasoning models (o1, DeepSeek-R1) โ they already do internal iteration
- Creative/open-ended generation โ repetition doesn't help much
- Latency-critical applications where token count matters
Sometimes we overcomplicate things. We chase RAG pipelines, fine-tuning, complex prompt chains.
And the solution is literally: say it twice.
๐ Link to the full paper in the first comment
๐ฉ If you want to integrate simple, auditable prompt strategies into your AI compliance stack, contact me at: dott.anghel.ai@gmail.com
#AI #PromptEngineering #LLM #AIAct #Compliance #AIGovernance #GoogleResearch #FRIA
7 months ago | [YT] | 7
View 1 reply
Dottor Anghel AI
๐ ๐ก๐๐จ๐ฅ๐ข๐ก๐ฃ๐๐๐๐: ๐ง๐๐ โ๐ช๐๐๐๐ฃ๐๐๐๐ ๐ข๐ ๐ก๐๐จ๐ฅ๐ข๐ก๐ฆโ ๐๐ข๐ฅ ๐๐๐ ๐ฆ
In modern LLMs we talk about โneuronsโ and โfeaturesโ, but most teams never see what they actually do.
Neuronpedia tries to fix that by turning neuron/SAE interpretability into a shared, navigable resource.
1๏ธโฃ ๐ช๐๐๐ง ๐ก๐๐จ๐ฅ๐ข๐ก๐ฃ๐๐๐๐ ๐๐ฆ
A public, collaborative atlas of neurons and SAE features for real models.
PRO โ
- Explanations, example prompts, and activations for individual neurons/features.
- Strong focus on sparse autoencoders (SAEs) and their interpretable โfeaturesโ.
- Web UI + APIs so you can browse, tag, and analyse features without building your own tooling.
2๏ธโฃ ๐ช๐๐๐ง ๐ฌ๐ข๐จ ๐๐๐ก ๐๐ข ๐ช๐๐ง๐ ๐๐ง
Think of Neuronpedia as your starting point for mechโinterp and steering.
PRO โ
- Inspect concrete features like โlegal toneโ, โEiffel Towerโ, or โselfโharmโ with real examples.
- Compare features across layers/models and plug SAEs into existing analysis libraries.
CON โ (๐ถ๐ณ ๐๐ผ๐ ๐ถ๐ด๐ป๐ผ๐ฟ๐ฒ ๐ถ๐)
- You keep treating steering vectors as blackโbox magic instead of grounded, documented concepts.
3๏ธโฃ ๐ฆ๐ง๐๐๐ฅ๐๐ก๐ & ๐ฆ๐๐๐๐ง๐ฌ ๐๐ก๐๐๐
Neuronpedia is also a bridge between interpretability and realโworld control.
- Use SAE features as steering knobs: upโ or downโweight a feature and see how outputs change.
- Design safety/style steering grounded in labeled features, not opaque directions, which reduces sideโeffects and helps with governance.
If you care about serious governance, this is a big upgrade over โwe changed the system prompt and it looks betterโ.
4๏ธโฃ ๐ช๐๐ฌ ๐ฃ๐ฅ๐๐๐ง๐๐ง๐๐ข๐ก๐๐ฅ๐ฆ ๐ฆ๐๐ข๐จ๐๐ ๐๐๐ฅ๐
- Cuts timeโtoโexperiment: hosting, visualisation, and collaboration are handled for you.
- Makes interventions easier to justify to stakeholders and regulators: you can point to specific, named features with examples, not just โprompt engineering that seems to workโ.
๐ฉ If you want to explore how Neuronpedia and steering vectors can fit into your AI governance or compliance stack, contact me at: dott.anghel.ai@gmail.com
#AI #Compliance #FRIA #AIAct #Neuronpedia
8 months ago | [YT] | 7
View 1 reply
Dottor Anghel AI
๐ ๐ ๐ข๐ฅ๐ ๐๐๐๐ง๐จ๐ฅ๐๐ฆ, ๐ ๐ข๐ฅ๐ ๐๐ข๐ก๐ง๐ฅ๐ข๐? ๐ก๐ข๐ง ๐ฅ๐๐๐๐๐ฌ. ๐ฆ๐๐ก๐๐๐โ๐๐๐๐ง๐จ๐ฅ๐ ๐ฆ๐ง๐๐๐ฅ๐๐ก๐ ๐ข๐๐ง๐๐ก ๐ช๐๐ก๐ฆ
In the SAE/steering world it sounds intuitive: โmore features = more controlโ.
Recent work suggests something else: more poorly chosen features can actually reduce coherence and control.
1๏ธโฃ ๐ฆ๐๐ก๐๐๐โ๐๐๐๐ง๐จ๐ฅ๐ ๐ฆ๐ง๐๐๐ฅ๐๐ก๐: ๐ฆ๐ก๐๐ฃ๐๐ฅ ๐ ๐ข๐๐
With an SAE, a single good feature is often enough to steer a behavior: refusal, legal tone, marketing style, etc.
โ
1. Clear, interpretable effect
2. Less interference with other behaviors
3. Easier to test, document, and defend in audits
โ (when you mix too many features)
1. Some mostly โreadโ the input instead of driving the output
2. Vectors add up and introduce noise โ text becomes less coherent, control less predictable
Moral: one wellโchosen feature > many โvibesโbasedโ features.
2๏ธโฃ ๐ ๐จ๐๐ง๐โ๐๐๐ฉ๐๐ ๐ฆ๐ง๐๐๐ฅ๐๐ก๐: ๐ฃ๐ข๐ง๐๐ก๐ง๐๐๐, ๐๐จ๐ง ๐ก๐ข๐ง ๐ ๐๐๐๐
Intervening on multiple layers / intermediate levels only helps if you know what you are doing.
โ
1. Some layers are better for content (middle), others for style/output (late)
2. Picking 1โ3 โkeyโ layers often beats โsteer everywhereโ
โ
There is no strong evidence that โapplying the same feature on many layersโ systematically beats using it where it has the most impact:
you risk more complexity, more noise, and governance explanations nobody really believes.
3๏ธโฃ ๐ง๐๐ ๐ฅ๐๐๐ ๐ง๐ฅ๐๐๐: ๐ฅ๐๐๐๐ง ๐ฉ๐๐๐ง๐ข๐ฅ + ๐ช๐๐๐โ๐ง๐จ๐ก๐๐ ๐๐๐๐ข๐๐๐ก๐
In practical tests (including EiffelโTowerโstyle demos), the quality jump does not come from โmore featuresโ but from:
1. Clamping
Limit how much the steering vector can distort activations โ the concept stays, with less overshooting and fewer weird outputs.
2. Lower temperature
Reduce sampling temperature โ less randomness, better instruction following, the steering signal is not washed out.
3. Light repetition penalty
Apply a moderate repetition penalty โ the concept remains present without obsessive repetition.
Net result: compared to โsimple addition + standard samplingโ, the combo single wellโchosen feature + clamping + lower T + light repetition penalty often yields outputs that are more coherent, better aligned, and still strongly express the target concept.
โ BONUS โ The same steering coefficient is not universal: its effect can swing a lot with different prompts and contexts. Recent work shows that for a fixed coefficient some inputs get strong steering, others weak or even inverted effects, so you must treat it as a promptโ and taskโdependent hyperparameter.
---
If you care about serious governance, the question is not โhow many features can I turn on?โ, but:
โWhich few features can I actually explain, control, and document โ and how do I set up decoding around them?โ
#AIAct #Compliance #Steering #Feature #AI #LegalTech
8 months ago | [YT] | 7
View 1 reply
Dottor Anghel AI
๐ ๐ฃ๐ฅ๐ข๐ ๐ฃ๐ง ๐๐ก๐๐๐ก๐๐๐ฅ๐๐ก๐ ๐ฉ๐ฆ ๐ฆ๐ง๐๐๐ฅ๐๐ก๐: ๐ช๐๐ข ๐ฆ๐๐ข๐จ๐๐ ๐โ๐๐๐ฉ๐๐๐ฆ ๐๐๐ง ๐ข๐ก?
Many companies are hiring prompt engineers.
Very few are building steering infrastructure.
If you need to govern models in production, this isnโt a โstyleโ choice โ itโs a choice about control, risk, and scalability.
1๏ธโฃ PROMPT ENGINEERING
โ PROS
1. No model access needed: the API is enough.
2. Fast iteration: you can tweak prompts in minutes.
3. Nonโtechnical teams can contribute.
โ CONS
1. Fragile: small changes in context can break behavior.
2. Hard to scale: each team โinventsโ its own prompts, so crossโproduct / crossโcountry consistency is weak.
3. Auditing is painful: itโs hard to explain to an auditor why a specific output appeared.
2๏ธโฃ STEERING
โ PROS
1. Internal control: you act on hidden states with concept vectors (risk, tone, safety, etc.).
2. Continuous adjustment: X_next = X + ฮฑ V โ you dial behavior up or down instead of rewriting prompts.
3. Central governance: you can define โcorporateโ vector libraries and reuse them across products and use cases.
โ CONS
1. Needs lowโlevel access (selfโhosted models or advanced tooling).
2. Technical setup: extracting, testing, and documenting robust vectors is not a side project.
3. Not every behavior is a โsingle directionโ: some vectors are less reliable and must be refined.
3๏ธโฃ HOW TO IMPROVE STEERING: FROM โJUST ADDโ TO FINEโGRAIN CONTROL
The basic scheme is:
- addition: X_next = X + ฮฑ V
- standard sampling: default temperature, no extra tuning.
Practical tests show you can make steering much more stable with:
1. Clamping
Limit how much the vector can deform activations (per dimension, norm, or token).
Effect: the concept stays, but you avoid overshooting and weird outputs.
2. Lower temperature
Reduce sampling temperature.
Effect: less randomness, stronger adherence to instructions, the steering vector isnโt washed out by noise.
3. Light repetition penalty
Apply a moderate repetition penalty.
Effect: the target concept remains present without obsessive, repetitive outputs.
Net result: compared to โaddition + standard samplingโ, the combo clamping + lower T + light repetition penalty produces smoother responses, better instructionโfollowing, and equal or better preservation of the target concept.
For prototypes and lowโrisk use cases, prompt engineering is enough.
For critical, auditable, multiโteam products, you need to move from โwriting promptsโ to designing real steering levers.
#AIAct #AI #PromptEngineering #Steering #Compliance
8 months ago | [YT] | 7
View 1 reply
Dottor Anghel AI
๐งฉ ๐ฆ๐ฃ๐๐ฅ๐ฆ๐ ๐๐จ๐ง๐ข๐๐ก๐๐ข๐๐๐ฅ๐ฆ (๐ฆ๐๐): ๐ง๐๐ ๐๐๐๐๐๐ก ๐ ๐๐ฃ ๐ฌ๐ข๐จ ๐ฆ๐๐ข๐จ๐๐ ๐๐ ๐จ๐ฆ๐๐ก๐ ๐ง๐ข ๐๐ข๐ฉ๐๐ฅ๐ก ๐ฌ๐ข๐จ๐ฅ ๐๐๐ ๐ฆ
Most companies control AI with prompts and policies.
SAEs let you control it with internal switches, instead of hoping the model โbehavesโ.
Steering tells you how to push a model.
Concept vectors tell you in which direction.
Sparse Autoencoders (SAE) tell you what is actually inside, in a readable way.
1๏ธโฃ ๐ช๐๐๐ง ๐๐ก ๐ฆ๐๐ ๐๐ข๐๐ฆ (๐ฆ๐๐ ๐ฃ๐๐๐๐๐๐)
An SAE takes an LLMโs activations (hidden states) and recodes them into:
1. a larger LATENT space
2. that is SPARSE: almost all values are zero.
Forcing sparsity has a key effect:
each active feature tends to represent a cleaner, more interpretable concept (or a small cluster of nearby concepts).
Instead of โpolysemanticโ neurons doing a bit of everything, you get a list of switches:
1. โlegaleseโ
2. โinformal languageโ
3. โhate speechโ
4. โcode / snippetsโ
etc.
In practice, the SAE turns the chaos of activations into a switchboard, where each button turns on a specific behavior of the model.
And because there are thousands of features, you can use LLMs themselves to help name each switch (โthis feature fires on phrases X: looks like โpolite toneโ, โlegaleseโ, etc.โ) โ this is autoโinterpretability.
2๏ธโฃ ๐๐ข๐ช ๐ฆ๐๐๐ฆ, ๐๐ข๐ก๐๐๐ฃ๐ง ๐ฉ๐๐๐ง๐ข๐ฅ๐ฆ & ๐ฆ๐ง๐๐๐ฅ๐๐ก๐ ๐๐๐ง ๐ง๐ข๐๐๐ง๐๐๐ฅ
Each sparse SAE feature has:
1. a vector in latent space (the โshapeโ of the concept)
2. an activation weight (how strongly it is on for that token / sequence).
This means you can:
1. Identify which features fire on certain concepts (toxicity, PII leaks, nonโcompliant tone).
2. Treat that feature as a concept vector: the direction in latent space representing that behavior.
From there you go back to the steering formula:
X_next = X + V * ฮฑ
- X = current internal state of the model
- V = vector of the SAE feature (the concept direction)
- ฮฑ = how much you amplify or damp that concept
Itโs the same mechanism as steering vectors: you press an internal button and push the modelโs behavior in that direction.
3๏ธโฃ ๐ช๐๐ฌ ๐๐๐๐๐ / ๐๐ง๐ข / ๐ฅ๐๐ฆ๐ ๐ฆ๐๐ข๐จ๐๐ ๐๐๐ฅ๐
For AI governance, SAEs matter because they let you:
1. Map โrisk zonesโ inside the model (features that light up on problematic content).
2. Build stable technical controls: not just prompts and policies, but mathematical levers over internal concepts.
Itโs the shift from:
โWe told the model not to do itโ
to
โWe directly lowered the activation of the features that cause that behaviorโ.
โก๏ธ ๐ก๐๐ซ๐ง ๐ฃ๐ข๐ฆ๐ง: ๐ฃ๐ฅ๐ข๐ ๐ฃ๐ง ๐๐ก๐๐๐ก๐๐๐ฅ๐๐ก๐ ๐ฉ๐ฆ ๐ฆ๐ง๐๐๐ฅ๐๐ก๐
Which one actually gives you control at scale? Weโll compare costs, risks, and tradeโoffs from a Cโlevel and controlโfunction perspective, not a โprompt hackerโ one.
#AIAct #AI #SAE
8 months ago (edited) | [YT] | 7
View 1 reply
Dottor Anghel AI
๐ง ๐๐ข๐ก๐๐๐ฃ๐ง ๐ฉ๐๐๐ง๐ข๐ฅ๐ฆ: ๐ช๐๐๐ก ๐ ๐๐ข๐ก๐๐๐ฃ๐ง ๐๐๐๐ข๐ ๐๐ฆ ๐ ๐ ๐๐ง๐๐๐ ๐๐ง๐๐๐๐ ๐๐๐ฅ๐๐๐ง๐๐ข๐ก
In previous posts we talked about:
- steering = changing an LLMโs behavior โon the flyโ
- concept embeddings & activation space = the internal map you steer on
Now comes the key piece: CONCEPT VECTORS.
A concept vector is a DIRECTION in activation space that corresponds to a specific concept:
โlegal toneโ, โmarketing styleโ, โtoxicityโ, โrisk conservatismโ.
๐ ๐๐ฟ๐ผ๐บ ๐ต๐ถ๐ฑ๐ฑ๐ฒ๐ป ๐๐๐ฎ๐๐ฒ ๐๐ผ โ๐๐๐ฒ๐ฒ๐ฟ๐ฒ๐ฑโ ๐๐๐ฎ๐๐ฒ
Imagine a hidden state vector of an LLM at some layer:
Hidden State Vector:
X = [1; 3; 6; 7]
Now suppose weโve found a concept vector:
V = [โ1; 0; 2; 1] (example)
When we do steering, we build a new state:
X_next = X + V * ๐ถ
where ฮฑ is a coefficient that controls how strongly we apply that concept (the โintensity knobโ).
In practice:
- X = how the model was โthinkingโ before
- V = the direction โmore legalโ, โless toxicโ, โmore conservativeโ
- ๐ถ = the dial: 0.2, 0.5, 1.5โฆ how hard you push the concept
Mathematically, itโs just vector addition.
Operationally, itโs a personality/behavior shift without touching the modelโs weights.
โ๏ธ ๐๐ข๐ช ๐ช๐ ๐๐๐ก๐ ๐ง๐๐ ๐ฉ๐๐๐ง๐ข๐ฅ ๐ฉ
Two main (simplified) paths:
1๏ธโฃ Difference between two prompt groups
- Group A: prompts that EXPRESS the concept (e.g. toxic answers).
- Group B: prompts that DO NOT express it (e.g. neutral answers).
You compute the mean activations for A and for B, then:
V โ mean_activations(A) โ mean_activations(B)
This V points in the direction โmore like A, less like Bโ.
2๏ธโฃ Sparse Autoencoders (SAE) โ teaser
Sparse Autoencoders take the modelโs activations and rewrite them into a larger but SPARSE space (mostly zeros).
Each sparse โfeatureโ often corresponds to an internal concept (or nearโconcept): an interpretable pattern of behavior.
In short:
- the SAE gives you a โbuttonโ (a feature) that turns a concept on/off
- that button corresponds to a vector in latent space you can use to steer the model
SAEs deserve their own post, so weโll go deeper next time.
๐งญ ๐ช๐๐ฌ ๐๐ข๐ก๐๐๐ฃ๐ง ๐ฉ๐๐๐ง๐ข๐ฅ๐ฆ ๐ ๐๐ง๐ง๐๐ฅ ๐๐ข๐ฅ ๐๐ข๐ฉ๐๐ฅ๐ก๐๐ก๐๐
They let you turn:
- โthis model is too aggressive/creativeโ
into
- โI apply +0.3 on prudence, โ0.2 on creativity in its internal activationsโ.
Itโs the shift from written policies to mathematical levers: controllable, measurable, reproducible.
โก๏ธ ๐ก๐๐ซ๐ง ๐ฃ๐ข๐ฆ๐ง: Sparse Autoencoders and how to use them to discover and control hidden concepts in your LLMs (without manual reverseโengineering).
๐ฉ Want help mapping concept vectors for risk, tone, or safety in your own AI systems โ and turning them into real governance levers, not just prompts? Write to me confidentially: dott.anghel.ai@gmail.com
#AI #AIAct #FRIA #ConceptVectors
8 months ago (edited) | [YT] | 7
View 1 reply
Load more