Accessible Agent Interfaces for Keyboard, Screen Reader, and Voice Users
Design streaming conversations, tool progress, approvals, errors, and voice controls that preserve focus, user control, and understandable status.
Stanley Yang turns complex workflows into clear, dependable full-stack products built for real-world demand.
Design streaming conversations, tool progress, approvals, errors, and voice controls that preserve focus, user control, and understandable status.
A migration playbook for adopting the evolving OpenTelemetry GenAI span, metric, and event schema safely.
A practical decision framework for choosing a multi-agent architecture, plus a minimal coordinator that keeps concurrency, budgets, and failure handling explicit.
Build a provider-neutral agent evaluation suite with deterministic checks, repeated trials, safety cases, uncertainty, and a CI policy that catches regressions.
Protect persistent agent memory with source trust, write gates, quarantine, contradiction checks, provenance, and complete recovery procedures.
Design tenant-safe agent memory with provenance, hybrid search, retention, deletion, and pgvector indexing that does not mistake every transcript for knowledge.
Make repeated agent prompts cacheable in vLLM without crossing tenant boundaries, and measure whether cache hits actually improve latency.
Design admission control, bounded queues, concurrency limits, deadlines, and graceful degradation for agent workloads.
A reproducible protocol for testing PostgreSQL 18 AIO across vector search, filtered retrieval, ingestion, vacuum, and cache states.
A reproducible protocol for comparing weight, activation, and KV-cache quantization without inventing savings or hiding quality regressions.
Measure speculative decoding against a target-only baseline across load, prompt types, acceptance rates, memory, latency, throughput, and output equivalence.
A reproducible protocol for measuring JSON validity, schema compliance, semantic accuracy, latency, and failure recovery across open models.
Enforce token, tool-call, latency, concurrency, and monetary ceilings across nested agent work without relying on prompt instructions.
Preserve page regions, reading order, tables, OCR alternatives, and source versions so every document answer can resolve to visible evidence.
Represent evidence at span level, require citations for verifiable claims, and abstain when retrieval cannot support a grounded answer.
Wrap any coding model in a reproducible harness that controls workspaces, commands, credentials, budgets, evidence, and verification.
Estimate weights, KV cache, runtime overhead, context concurrency, and replica demand—then replace assumptions with measured service curves.
Route remote MCP traffic by authenticated tenant while isolating tokens, sessions, discovery, quotas, results, and audit records.
Implement an agent loop with typed plans, evidence-based review, bounded revision, explicit stop states, and safe tool execution.
Design a restartable RAG ingestion pipeline that preserves provenance, detects changes, avoids duplicate work, and removes stale chunks from retrieval.
Design a reconnectable browser voice agent with explicit signaling, ephemeral authorization, audio lifecycle, interruption, and observable session state.
Build and test a real MCP stdio server, then make the transport, validation, error, and authorization decisions required for a remote deployment.
Create resettable, observable email, calendar, browser, and database environments that test agent effects without touching real systems.
Reserve one GPU for generation, run embeddings and retrieval on CPU, store cited chunks in Postgres with pgvector, and evaluate retrieval before answers.
Create resettable visual web tasks with programmatic outcome validators, layout perturbations, safety checks, and reproducible multimodal traces.
Define typed WIT contracts, explicit capabilities, resource limits, provenance, and conformance tests for portable agent tools.
Design a review agent that reports reproducible defects with precise evidence instead of producing noisy summaries and style opinions.
Design an MCP tool that serves an interactive UI resource, receives structured results, and stays safe across hosts with different extension support.
Connect to a local MCP server, verify capability discovery, call a typed tool, and handle remote transports without confusing discovery with authorization.
Turn the OWASP Top 10 for Agentic Applications into safe, repeatable abuse cases with simulated tools, hard invariants, evidence, and CI gates.
Capture actor chains, policy decisions, approvals, tool effects, and evidence so investigators can reconstruct an agent incident without logging hidden reasoning or secrets.
Design a durable MCP tool call with polling, deferred results, input-required handling, cancellation, and tenant-safe retention.
Stream agent UI and typed run events with Suspense, Route Handlers, cancellation, backpressure, and accessible status updates.
Connect a coordinator to an independently operated report agent using an Agent Card, SendMessage, task state, typed artifacts, and strict authorization.
A correctness-first guide to caching model prefixes, retrieval results, tool calls, and final AI outputs safely.
Replace broad ambient credentials with narrow, expiring, attenuable grants bound to one agent task, tenant, operation, and resource set.
Select open weights with a reproducible, product-specific suite covering capability, tool use, safety, latency, memory, and operational fit.
Choose between APIs, structured browser automation, accessibility trees, and visual computer use using reliability, coverage, security, and cost.
Design a context pipeline that selects trusted instructions, relevant evidence, tool definitions, and recent state within an explicit token budget.
Test agent tool schemas, transport behavior, authorization, retries, and compatibility with valid, invalid, and adversarial requests.
A practical ledger for attributing model, tool, compute, retry, and shared costs to successful multi-agent outcomes.
Preserve user intent and actor identity across agent-to-service and agent-to-agent calls with narrow audiences, scopes, consent, and delegation chains.
Use ratified WASI 0.3 async components with pinned runtime support, explicit capabilities, deadlines, backpressure, and safe cancellation.
Turn human approval into a real security control by previewing exact effects, binding approval to immutable action data, and preventing replay or substitution.
Turn design states into deterministic Playwright checks that combine semantic assertions, visual baselines, accessibility, and human-reviewed diffs.
Record model, tool, time, randomness, configuration, and state dependencies so agent failures can be replayed without repeating real-world effects.
Separate prefill and decoding only when independent scaling and latency isolation beat KV-transfer cost, duplicated weights, and operational complexity.
Measure absent, concise human-written, generated, and noisy repository instructions with paired tasks, fixed agent settings, and no invented results.
Design long-running agent workflows that survive crashes, retry only safe work, bind approvals to exact actions, and reconcile ambiguous side effects.
Choose AI backend placement with a reproducible latency, quality, capacity, cost, security, and failure benchmark.