Modern ecosystems are changing faster than humans can observe. Field researchers rely on slow, manual measurements, while satellites offer delayed, low-resolution snapshots. There’s a massive sensing gap between the ground and the sky.

My mission: build the robotics stack that makes ecological monitoring cheaper, longer, repairable, and more intelligent; drones, sensors, and edge-AI systems for real-time environmental understanding.

I design the entire stack myself:
C-based transformers → HPC training pipelines → embedded flight controllers → sensors → ecological applications.
Every prototype I publish — BLDC controllers, CPU-optimized LLMs, IMU calibration rigs, test benches; is a building block toward the world’s most accessible ecological intelligence system.

If you’re interested in:
Embedded AI on CPUs
Drones, flight dynamics, attitude control
C/MPI/HPC engineering
Real-world conservation through robotics
Deep technical builds from first principles

This channel is for you.



ANTSHIV ROBOTICS

Three audio models. Two CPU systems. The same 42:22 recording. No GPU.

C-Kernel-Engine now completes Whisper Base, NVIDIA Parakeet TDT 0.6B v3, and Cohere Transcribe 2B through generated C model circuits. I ran the same source audio at the same CKE revision on my Intel P3 and AMD Ryzen nodes.

Whisper Base
P3: 121.74 seconds, 20.88x real-time
Ryzen: 65.25 seconds, 38.96x real-time

Parakeet TDT 0.6B v3
P3: 1,515.19 seconds, 1.68x real-time
Ryzen: 881.95 seconds, 2.88x real-time

Cohere Transcribe 2B
P3: 1,280.99 seconds, 1.98x real-time
Ryzen: 583.02 seconds, 4.36x real-time

The transcript artifact for each model matched exactly between Intel AVX2 and AMD AVX-512. That demonstrates deterministic cross-host execution for these retained runs. It does not claim that the three models have equal transcription quality. Cohere also uses a pinned external speech schedule and infers 2,288.04 seconds of the recording, so its timing policy differs from the other paths.

Whisper is practical in my production workflow today. Parakeet and Cohere now work end to end through CKE's generated code, and the measurements expose where CPU utilization and kernel performance still need work.

Read the complete article:
www.shivasnotes.com/blog/5981/from-native-kernels-…

Watch the exact 42:22 source video used for the benchmark:
https://youtu.be/XMyPVVmjuiY?si=gEwVi...

Watch the latest follow-on video about the networking ideas that influenced CKE:
https://youtu.be/cHUGuZs4htw?si=KFxu1...

CKE hardware lab and the roles of the P3 and Ryzen nodes:
c-kernel-engine.github.io/C-Kernel-Engine/support-…

C-Kernel-Engine:
github.com/C-Kernel-Engine/C-Kernel-Engine

#CKE #CPUAi #Whisper #Parakeet #Cohere #SpeechRecognition #KernelEngineering

6 days ago | [YT] | 1

ANTSHIV ROBOTICS

How do you get started with C-Kernel-Engine (CKE)?

This carousel follows the smallest useful path: prepare Linux, clone CKE and its submodules, run a small Gemma 3 model on a CPU, inspect the generated C pipeline, and understand what one coherent result does and does not prove.

The easiest prerequisite is an older x86-64 laptop or desktop you can repurpose as a native Linux node. This keeps the first experiment inexpensive, gives you direct access to compilers, perf, CPU affinity, NUMA and memory controls, and turns learning Linux into a useful side quest. It also lets you build and inspect AI software on compute you own instead of making every experiment depend on a rented platform.

CKE combines explicit model circuits, kernel capability maps, memory planning, generated C, native CPU execution, and numerical evidence. It is still an emerging research system, so the artifact, quantization, context length, CKE commit, processor, and tested boundary all matter.

Read the complete getting-started guide:
www.shivasnotes.com/blog/5979/how-to-get-started-w…

New to the architecture? Start with What Is the C-Kernel-Engine?:
www.shivasnotes.com/blog/5889/what-is-the-c-kernel…

C-Kernel-Engine repository:
github.com/C-Kernel-Engine/C-Kernel-Engine

Official CKE quickstart:
c-kernel-engine.github.io/C-Kernel-Engine/quicksta…

CKE concepts:
c-kernel-engine.github.io/C-Kernel-Engine/concepts…

CKE v8 runbook:
c-kernel-engine.github.io/C-Kernel-Engine/v8-runbo…

Model and kernel evidence matrix:
c-kernel-engine.github.io/C-Kernel-Engine/model-ke…

#CKE #CPUAi #KernelEngineering #OpenSourceAI #Linux

1 week ago | [YT] | 8

ANTSHIV ROBOTICS

Three audio model families. One 42-minute recording. CPU only.

C-Kernel-Engine now runs OpenAI Whisper, NVIDIA Parakeet TDT 0.6B v3 and Cohere Transcribe across my complete 42:22 kernel-engineering presentation on owned CPU compute.

On the Intel Core i7-14700T P3:

Whisper Base: 2:24.24, or 17.62x real-time.
Cohere Transcribe 2B: 26:53.7, or 1.58x real-time.
Parakeet TDT 0.6B v3: 35:33.1, or 1.19x real-time.

Whisper is already practical in my video workflow. Parakeet works deterministically across Intel and AMD, but 94.9% of its current runtime is in the FastConformer encoder and CPU utilization remains low. Cohere now emits bounded, resumable JSON, text and 729 SRT captions across 123 speech slices. Its complete selected-token, caption, transcript and SRT outputs match exactly between Intel and AMD.

This is not a perfectly symmetric model benchmark. Cohere uses a pinned external Silero VAD schedule and intentionally skips 254 seconds classified as silence. The three architectures and long-audio policies differ. CrispASR agreement measures implementation parity, not human transcription accuracy. VAD ownership, forced-alignment accuracy, multilingual testing, quantization, concurrency, diarization and further CPU optimization remain open.

Read the complete article:
www.shivasnotes.com/blog/5978/beyond-whisper-cke-r…

Watch the complete 42-minute source video transcribed in these tests:
https://youtu.be/XMyPVVmjuiY

C-Kernel-Engine:
github.com/C-Kernel-Engine/C-Kernel-Engine

Cohere long-audio PR #510:
github.com/C-Kernel-Engine/C-Kernel-Engine/pull/51…

Parakeet long-audio PR #507:
github.com/C-Kernel-Engine/C-Kernel-Engine/pull/50…

CKE Cohere audio documentation:
c-kernel-engine.github.io/C-Kernel-Engine/v8-coher…

CKE Parakeet TDT circuit and evidence:
c-kernel-engine.github.io/C-Kernel-Engine/v8-parak…

CKE audio-kernel deep dive:
c-kernel-engine.github.io/C-Kernel-Engine/audio-ke…

CKE kernel catalogue:
c-kernel-engine.github.io/C-Kernel-Engine/kernels.…

#CKE #CPUAi #Whisper #Parakeet #Cohere #SpeechRecognition #KernelEngineering

2 weeks ago | [YT] | 1

ANTSHIV ROBOTICS

Qwen3.8 Flash Next did not fit C-Kernel-Engine perfectly.

The model introduced a combination CKE had never represented as one complete circuit: four hyper-connection streams, Per-Layer Embedding, recurrent Gated DeltaNet, query-selected sparse attention, 512 routed experts and a gated shared expert. Supporting it required new kernels, kernel maps, circuit vocabulary, memory contracts and some central compiler work.

This carousel presents that work honestly. The bring-up changed 127 files, including five core compiler or planner files. Some changes exposed breakage before they hardened CKE: invalid generated identifiers, ambiguous provider fallback, stale X-Ray runtime bundles, mislabeled prefill capability and a shared-expert scheduling difference that became visible after quantization.

The repair direction is contract-first. Known model values belong in circuit and kernel-map JSON. A genuinely new semantic may require one generic DSL extension, but later models should reuse that contract without model-family conditionals. Producer and consumer memory edges must declare their shape, dtype, lifetime, persistence and aliasing rules rather than relying on guesses.

The tested Q4_K_M text path now has bit-exact short-prompt inference evidence. That does not mean full BF16, multimodal, long-context or training certification. Inference is the first step: it proves the forward circuit and kernels. CKE's longer objective is to add explicit backward kernels, gradients, training-state ownership and memory plans so kernels can eventually be composed and trained on owned CPU compute.

Read the full article:
www.shivasnotes.com/blog/5976/how-cke-brought-up-q…

Review the implementation and evidence:
github.com/C-Kernel-Engine/C-Kernel-Engine/pull/45…

CKE kernel concepts:
c-kernel-engine.github.io/C-Kernel-Engine/concepts…

CKE kernel maps:
c-kernel-engine.github.io/C-Kernel-Engine/kernel-m…

Qwen3.8 Flash Next runbook:
c-kernel-engine.github.io/C-Kernel-Engine/v8-runbo…

Want to contribute CPU kernels, compiler contracts or numerical fixtures?
github.com/C-Kernel-Engine/C-Kernel-Engine/blob/ma…
discord.gg/MZUWqGaVb

Keywords: C-Kernel-Engine, CKE, Qwen3.8 Flash Next, CPU inference, compiler architecture, kernel maps, DSL, numerical parity, backpropagation, CPU training

#CKE #CPUAi #KernelEngineering #Qwen

3 weeks ago | [YT] | 5

ANTSHIV ROBOTICS

I used C-Kernel-Engine's Whisper support to transcribe and help edit my own kernel-engineering video. The first result was broken. That failure led us from a graph-wiring regression to a second 40 ms audio-tail bug, and finally to a permanent five-minute nightly test covering five Whisper model sizes.

The carousel explains what CKE did, what FFmpeg and Kdenlive did, why the generated encoder diverged, how the repair was measured, and why a successful model demonstration must become a retained end-to-end regression fixture.

Most software I build needs to become useful in my own work. I have developed Antsand for roughly 14 years, and it now produces my websites, ShivasNotes, and much of my content workflow. CKE is beginning the same transition from research system to practical tool. Whisper transcription is the clearest first application. Over the next 4–6 months, as I add model families, compute, and stronger test harnesses, I expect CKE to become an integral part of my AI workflow and eventually connect more deeply with Antsand and Antshiv Robotics development.

Many runtimes can run OpenAI's open Whisper model. Running it is evidence that CKE works, but it is not CKE's unique north star. The longer-term objective is to make forward kernels, backward kernels, dtypes, quantization formats, and model circuits reusable enough to compose and eventually train models on owned CPU compute.

Watch the finished video:
https://youtu.be/XMyPVVmjuiY

Read the full Whisper regression investigation:
www.shivasnotes.com/blog/5974/i-used-cke-to-transc…

Read the related Qwen3.8 regression investigation:
www.shivasnotes.com/blog/5975/qwen3-8-broke-in-cke…

C-Kernel-Engine repository:
github.com/C-Kernel-Engine/C-Kernel-Engine

CKE kernel concepts:
c-kernel-engine.github.io/C-Kernel-Engine/concepts…

Antsand, the content management system I have been building since 2012, still powers my websites, blogs, content marketing, and broader content strategy:
www.antsand.ca/

Keywords: C-Kernel-Engine, CKE, Whisper, CPU AI, speech recognition, compiler regression, numerical testing, nightly testing, video editing, FFmpeg, Kdenlive

#CKE #Whisper #CPUAi #KernelEngineering

3 weeks ago (edited) | [YT] | 4

ANTSHIV ROBOTICS

How Audio Becomes Words: Whisper's Kernels On A CPU

Everyone treats speech-to-text as one magic model. It isn't. Whisper is a chain of small, inspectable kernels — and once you can see each one, the whole "how does audio become text?" mystery falls apart. The real engineering question is: can you run that entire chain as plain, readable C on a CPU, with no GPU, and still match the reference token for token?

This carousel walks the whole path Whisper takes inside C-Kernel-Engine (CKE): raw sound → mel spectrogram → encoder → decoder → transcript. You will see why 30 seconds of audio always collapses to 1500 tokens, why the encoder uses full unmasked attention while the decoder is causal, and how cross-attention into a frozen audio memory makes decoding roughly 79× cheaper per step. You will also see how tiny, medium and large Whisper are literally the same stitched circuit at different depths.

Read the full ShivasNotes post:
www.shivasnotes.com/blog/5973/how-audio-becomes-wo…

What this covers:
- The 8 frontend kernels that turn sound into a tensor (only the Conv1D stem is learned)
- Why 480,000 samples → 3000 STFT frames → 1500 audio tokens
- The encoder: full 1500×1500 attention, no mask, no causality
- The decoder: causal self-attention plus cross-attention into cached audio memory
- The full stitched circuit: one block stamped 4 / 24 / 32 times
- Tiny, base, small, large — what actually changes as you scale
- Why generated FP32 C stays token-exact vs Hugging Face on a CPU

Related:
- C-Kernel-Engine audio kernels deep dive: c-kernel-engine.github.io/C-Kernel-Engine/audio-ke…
- C-Kernel-Engine: github.com/c-kernel-engine/C-Kernel-Engine

Keywords:
Whisper, speech to text, ASR, audio kernels, mel spectrogram, STFT, log-mel, encoder decoder, cross attention, KV cache, CPU inference, C-Kernel-Engine, FP32 C, token-exact

#Whisper #AIKernels #CPUInference #CKernelEngine

3 weeks ago | [YT] | 2

ANTSHIV ROBOTICS

Antshiv Robotics is currently hardening two large engineering programs.

The first is C-Kernel-Engine: an auditable CPU-native AI runtime that describes modern models as explicit circuits, resolves legal kernels, plans memory, generates native C, and compares the result with PyTorch and llama.cpp.

The second is the Antshiv flight-control stack: aircraft contracts, quaternion and state-estimation mathematics, controller tuning, geometry-derived mixing, checked rigid-body simulation, nRF5340 firmware, CEVA sensor integration, and a gradual path from desktop SIL to HIL and physical flight.

They are different systems, but the engineering method is similar:

- Keep the important mathematics and ownership explicit.
- Build portable C boundaries rather than hiding everything inside a framework.
- Compare against independent references.
- Make incorrect states fail rather than quietly continue.
- Publish the commands, measurements, limitations and remaining work.

The programs may eventually connect, but not by making AI responsible for basic flight safety. Stabilization, navigation, failsafes and recovery must remain deterministic and independently functional. CKE can later add optional onboard perception and larger portable ground-compute capabilities.

Today’s flight-control progress report:
www.shivasnotes.com/blog/5972/today-i-made-my-dron…

C-Kernel-Engine:
github.com/C-Kernel-Engine/C-Kernel-Engine

AeroDynControlRig:
github.com/antshiv/AeroDynControlRig

Antshiv Robotics Flight Controller:
github.com/antshiv/ASR-FC

Join the Antshiv Robotics Discord:
discord.gg/MZUWqGaVb

#AntshivRobotics #CPUAI #FlightController #Robotics #ControlSystems #EmbeddedSystems #CKE

1 month ago | [YT] | 4

ANTSHIV ROBOTICS

CKE is slowly becoming easier to extend.

Over the last week, C-Kernel-Engine added or hardened CPU execution support for AMD Instella-MoE, Cohere Command R, Poolside Laguna-XS 2.1, NVIDIA Nemotron, Moonshot AI's Kimi-VL text decoder, and Qwen3.8.

This carousel explains what that actually means. CKE inspects a checkpoint, selects a model-owned circuit, resolves legal kernels through kernel maps, plans memory, generates native C, and checks the result against PyTorch or llama.cpp. The model families share a compiler path, but they do not share one generic transformer graph.

Each family pressured a different part of the architecture: Cohere's parallel residual path, Laguna's global/sliding attention and softplus head gate, Instella's BF16 gated MLA and FarSkip streams, Nemotron's Mamba2 state and heterogeneous layer policy, Kimi's MLA decoder, and Qwen3.8's dense DeltaNet/full-attention circuit.

The evidence is intentionally separated. Some lanes currently prove coherent end-to-end execution. Others have sampled numerical agreement. Qwen3.8 has a 4,096-position trajectory-exact result at the tested llama.cpp boundary. Performance is tracked separately from correctness.

Full article:
www.shivasnotes.com/blog/5971/cke-now-runs-more-th…

CKE documentation and model-kernel matrix:
c-kernel-engine.github.io/C-Kernel-Engine/model-ke…

CKE source:
github.com/C-Kernel-Engine/C-Kernel-Engine

Join the Antshiv Robotics Discord:
discord.gg/MZUWqGaVb

Keywords: C-Kernel-Engine, CKE, CPU AI, generated C, model circuits, kernel maps, numerical parity, Qwen3.8, Instella, Cohere Command R, Laguna, Nemotron, Kimi-VL, llama.cpp, PyTorch

#CPUAI #OpenModels #KernelEngineering #CKE

1 month ago | [YT] | 1

ANTSHIV ROBOTICS

Qwen3.8 is punching way above its weight class.

I put its Ryzen-generated infographics beside comparable Opus work. Nobody looking at the finished artifacts could reliably tell which system made which one.

Each run first prefilled exactly 131,072 input tokens: a large CKE dossier the model had to synthesize. Both averaged about 32.6 prefill tok/s. The outputs were 10,546 tokens in 3 h 12 m and 12,472 tokens in 3 h 37 m.

I ran it locally, went to sleep, and woke up to several infographics. Three larger diagrams are in the article, and they honestly look amazing. It was slower, but the reasoning, synthesis and artifact quality matched the capability I got from the frontier model.

As CKE tests distributed CPU execution and stronger nodes, these times should fall. If models in this capability class keep being released and become smaller, most of my practical work can move to owned hardware. In the coming months, or within a year or two, I may no longer need ChatGPT or Claude subscriptions for most of what I do.

Read the full article and generated Qwen3.8 SVG at:
www.shivasnotes.com/blog/5970/one-ryzen-128k-conte…

1 month ago | [YT] | 51

ANTSHIV ROBOTICS

CKE added Qwen3.8-27B in about an hour. The interesting part is not the clock by itself: the model did not require a new arithmetic kernel, a model-specific runtime fork, or a new compiler path. CKE needed to identify the artifact honestly, define a model-owned dense circuit, reuse its existing numerical providers, generate native C, and certify the result against llama.cpp.

Read the complete article and evidence:
www.shivasnotes.com/blog/5969/cke-added-qwen3-8-in…

The first native run used the standard 17.8 GB Qwen3.8-27B Q4_K_M artifact on CKE's AMD Ryzen 9 9950X3D node. It mapped all 851 required weights and validated 1,741 prefill operations plus 1,628 decode operations. PR #388 then matched all 4,000 top-1 trajectory rows and roughly 993 million vocabulary values bit-for-bit against pinned llama.cpp. X-Ray added 4,011 exact internal rows across recurrent and full-attention sentinel layers.

This carousel also explains the larger objective. Model compatibility is evidence, not CKE's final destination. Open weights and technical reports let a small independent project learn from research funded at a scale it cannot reproduce. CKE is trying to turn those kernels, dtypes, quantization formats, circuits, gradients, and ISA providers into reusable building blocks for inference, model composition, and eventually training on commodity CPU hardware.

CKE documentation:
c-kernel-engine.github.io/C-Kernel-Engine/

CKE source:
github.com/C-Kernel-Engine/C-Kernel-Engine

Qwen3.8 bring-up PR #387:
github.com/C-Kernel-Engine/C-Kernel-Engine/pull/38…

Long-trajectory certification PR #388:
github.com/C-Kernel-Engine/C-Kernel-Engine/pull/38…

The founder-funded CKE CPU lab:
www.shivasnotes.com/blog/5967/i-spent-cad-9-575-bu…

Keywords: C-Kernel-Engine, CKE, Qwen3.8, Qwen3.8-27B, CPU AI, generated C, model circuits, kernel compiler, numerical parity, X-Ray, llama.cpp, AMD Ryzen 9 9950X3D, AVX-512, Q4_K_M, open weights, AI training on CPUs

#CKE #Qwen38 #CPUAI #OpenSourceAI

1 month ago | [YT] | 2