Blog tag

#Performance

31 posts tagged with Performance.

← Back to all posts
2 min read

Twelve Iterations on a Contact Map and a Snappier Shared Inbox

A day of visual iteration on the contact map, plus performance and correctness work in the shared inbox.

FrontendDesignPerformanceEmailSEA
Read more
2 min read

Pre-Integration Guards, Faster Mail, and the Sales Shared Inbox

A 46-commit day making the integration safe to deploy, speeding up Mail, and building the sales shared-inbox surface.

LaravelPerformanceEmailIntegrationDeployment
Read more
2 min read

Candle Continuity, Local-First Market Data, and Deploy Health

A 32-commit day making candle-service fast and continuous while moving dex-trader to local-first data.

FastAPISQLitePerformanceTradingData Pipelines
Read more
2 min read

Regime Feeds, Volume Workers, and Curation in Candle Service

Adding a Binance regime feed, backfilling volume, and dropping a concept that didn't earn its complexity.

PythonMarket DataSQLitePerformanceRefactor
Read more
2 min read

Building Edge Lab and Porting It to a Triple-Arc XPU Rig

Phase 1 through 9 of a trading research platform, then a hard pivot from numba-dpex to torch.xpu, Triton, and IPEX.

PythonTritonIntel ArcPerformanceBacktesting
Read more
2 min read

OpCache, JIT, and a Slimmer Docker Build

Trimming image build time while turning on PHP performance features and extracting a reusable migration kit.

DockerPHPPerformanceOpCacheSQLite
Read more
2 min read

26 Iterations: Porting Cleanroom Kernels to Gemma 4

A single day of rapid iteration took the Gemma 4 port from a baseline to validated QKV fusion, invalidating several assumptions along the way.

SYCLKernelsLocal LLMPerformanceIteration
Read more
2 min read

All-Local 29 Tokens Per Second and Fused Norm Kernels

Fusing QKV, inlining norm-weight application, and pushing block loads through the Q4_K path.

SYCLKernelsPerformanceLocal LLMOptimization
Read more
2 min read

Standalone Inference Working: Shared Memory and a CPU Fallback

The standalone backend finally produced inference after fixing ABI layout, shared-memory buffers, and a simple CPU fallback.

SYCLKernelsABILocal LLMPerformance
Read more
2 min read

Enabling Local Dispatch Op by Op Until the Graph Ran Entirely on My Kernel

47 commits and a progressive enablement series that ended with full local decode dispatch at 7.0 t/s and 100% correctness.

SYCLKernelsPerformanceDebuggingLocal LLM
Read more
3 min read

Fixing 20-60 Second Page Loads Across the AustinsElite Codebase

A performance day: SQLite query fixes, frontend bottlenecks, and performance smoke tests to keep it from regressing.

SQLitePerformanceLaravelPlaywrightFrontend
Read more
2 min read

Writing an ESIMD Fused Kernel From Reverse-Engineered Assembly

SLM-tiled GEMM, SPIR-V decompilation, and a corrected Q4_K dequant that finally produced coherent output.

SYCLESIMDKernelsReverse EngineeringPerformance
Read more