Deep-dive technical education for software engineers. Master retrieval, memory and distributed caching for production LLM applications through structured digital products you download the moment you buy.

A video course on caching, memory and state for LLM applications: cluster configuration, persistence strategies and failover for production workloads.

The complete written guide to embeddings, hybrid retrieval and building retrieval-augmented generation pipelines that hold up under real traffic.

Prompt patterns, evaluation checklists and Redis command references for engineers shipping AI features every week.

Start from zero: keys, expiry, data types and the mental model that makes every later Redis decision obvious.

Build durable event pipelines with Streams: consumer groups, acknowledgement, replay and backpressure for inference workloads.

Cut token spend by serving answers you already paid for: similarity thresholds, invalidation and safety rails for cached generations.

Build the test harness your retrieval pipeline is missing: golden sets, offline scoring and regression gates in CI.

Tool calling, planning loops, state machines and the failure modes that make autonomous agents unusable in production.

Short-term context, episodic recall and long-term profiles: designing memory that improves answers instead of bloating prompts.

Version, test and review prompts like code: templates, fixtures, diffs and rollback for teams past the copy-paste stage.

Traces, token accounting, latency budgets and quality signals — the dashboards you need before your first bad week.

Routing, batching, distillation and caching: a systematic method for halving model spend without hurting output quality.

Sharding, resharding, upgrades and incident response for clusters that carry production AI traffic.

Prompt injection, data exfiltration, tool abuse and tenant isolation — threat modelling for systems that take instructions from text.

Lawful basis, retention, deletion and vendor due diligence for teams putting personal data anywhere near a model.

Ingestion, deduplication, incremental reindexing and backfills for corpora that change every day.

Fuse BM25 and vector scores, apply business rules and tune ranking with feedback you already collect.

Server-sent events, partial rendering, cancellation and error recovery for interfaces that respond token by token.

When fine-tuning beats prompting, how to build the dataset, and how to evaluate the result before it reaches users.

Namespacing, quotas, noisy-neighbour control and per-tenant cost attribution for AI features sold to many customers.

Access-pattern-first schema design: key naming, secondary indexes, denormalisation and the trade-offs behind each choice.

Tail latency, timeouts, hedging and budget allocation across retrieval, model and tool calls.

Feature flags, staged rollout, human review queues and kill switches for behaviour you can't fully predict.
Download links are issued immediately after a successful checkout and emailed to you. Nothing is shipped physically.
Every course includes the full source code, deployment scripts and architectural diagrams used in the lessons.
If a product doesn't help your work, request a refund within 14 days of purchase and we'll return the full amount.
“Finally an AI resource that goes past demo notebooks. The architecture course helped us fix our retrieval cache and cut inference spend substantially.”
“The reference pack is always open on my second monitor. The most useful technical asset I've bought this year for my platform team.”
After checkout you are redirected to a secure download page, and an email with your permanent download links is sent to the address you used at checkout.
All prices are in pounds sterling (GBP) and include any tax that applies to your location. The total is confirmed before you pay.
Yes. For teams of five or more we offer discounted multi-user licences. Contact us for a quote.
Products are revised as the underlying tools change. Existing customers receive those updates at no extra cost.