Posts
All the articles I've posted.
-
Capacity Planning Qwen3-TTS on an A10G
A single-GPU capacity study for streaming Qwen3-TTS: SLOs, vLLM-Omni comparison, fleet estimates, and a proposed production architecture.
-
How to Build a Low-Latency Streaming Qwen3-TTS Server
Turning Qwen3-TTS into a deployable AWS streaming service with CUDA graphs, FastAPI WebSockets, and measured sub-200ms first-audio latency.
-
Advanced Retrieval for Retrieval-Augmented Generation
Query expansion, cross-encoder re-ranking, and embedding adaptors for improving RAG retrieval quality.
-
LLMs Evals: A General Framework for Custom Evaluations
A general framework for building rule-based and model-graded evaluations for LLM-based applications.