EP228: How SSH Works

Get a state-of-the-art, fully assembled agent harness. (Sponsored)

Strands harness is the open-source, state-of-the-art, fully assembled agent harness. It’s a complete, production-ready agent the moment you install it. Across six benchmarks it matched or beat other harnesses on accuracy while using 28% fewer tokens.

You get shell, file, and web tools, plus context management, memory, and subagents, all tuned and ready to use. Every default is yours to change. Run it as a library or from the strands CLI, on any model and any cloud. It’s built on the open-source Harness SDK, so you can extend or replace anything.

Get started


This week’s system design refresher:


How SSH Works

Secure shell is a method to access a remote machine securely over an unsecured network. It starts with a TCP connection from the SSH client to the remote machine (SSH server).

Both the machines exchange the SSH versions and negotiate crypto algorithms, and each side runs a key exchange protocol. The SSH server sends the public key and a signature, which the client machine verifies. That host key is then checked against the known_hosts file.

Both the machines then derive session keys on their end and do not share them over the network. Same keys, computed separately, never put on the wire. Then the SSH client starts authentication by sending the public key for login.

The SSH server matches the public key in the authorized_keys, and the SSH client also signs the auth request with the private key and sends the digital signature to the SSH server.

The private key never leaves your machine. The server will verify the signature with the client public key. This completes the SSH authentication, and now the session is open for communication.


Top 6 Techniques to Make Your AI System Efficient

Serving LLMs at scale is expensive. These 6 popular techniques reduce latency and cost, and make your system more efficient.

  1. Streaming: The model sends each token as soon as it becomes ready. This changes the user’s perceived latency to time-to-first-token.
  2. Quantization: Convert the model's weights to lower precision like FP8. This conversion leads to less memory usage, which translates to faster and cheaper serving.
  3. Continuous batching: With this technique, new requests are added to the batch once a slot becomes available. Therefore, the GPU is less idle.
  4. Prefix caching: This technique caches internal calculations for given input prompts. At runtime, if the prompt has a cached prefix, those calculations are skipped.
  5. Paged KV cache: This technique stores the cache in fixed-size blocks instead of one contiguous chunk. At runtime, blocks are allocated as the sequence grows, so memory is not reserved up front.
  6. Speculative decoding: Here, a small model (speculator) proposes multiple tokens. After that, the main model verifies them in one forward pass.

Over to you: what is missing from this list?


[Webinar] How to stop babysitting your agents (Sponsored)

Agents can generate code. Getting it right for your system, team conventions, and past decisions is the hard part. You end up wasting time and tokens in the correction loops.

More MCPs, rules, and bigger context windows give agents access to information, but not understanding. The teams pulling ahead have a context layer to give agents exactly what they need for the task at hand.

Join us for a FREE webinar on Oct 7 to see:

  • Where teams get stuck on the AI maturity curve and why common fixes fall short
  • How a context layer solves for quality, efficiency, and cost
  • Live demo: the same coding task with and without a context layer

If you want to maximize the value you get from AI agents, this one is worth your time.

Register Now


The Anatomy of an AI Agent

An AI agent can be thought of as a simple While-loop.

It uses an LLM to select an action, executes that action, evaluates the result, and repeats the process until the task is complete. Let’s take a closer look at each of these components:

  • Brain: The LLM is the core. It reads the situation, thinks, and decides what to do next. The big shift from chatbot to agent: the model isn’t writing text anymore, it’s making choices.
  • Planning: Hard tasks need more than one step. Agents break them down using methods like Chain of Thought (think step by step), Tree of Thoughts (try options, pick the best), or
    Reflexion (learn from mistakes and retry). Planning turns a fuzzy goal into clear actions.
  • Tools: An LLM without tools is a brain in a jar. Tools are functions the model can call, like web search, code execution, APIs, files, or browsers (often using the MCP standard). The model requests a tool, the system runs it, and the result comes back.
  • Memory: Without memory, every turn starts from zero. Short-term memory is the context window. Long-term memory lives in vector stores, files, and knowledge bases. When the window fills up, agents summarize old turns and carry the summary forward.
  • Loop: All four pieces work together in a cycle. The agent looks at the current state, decides what to do, uses a tool, sees the result, and repeats. It keeps going until it gives a final answer.
  • Guardrails: Not strictly anatomy, but important. Sandboxing, human checks, token limits, output validation, and scope limits keep autonomy from turning into expensive chaos. The more autonomy you give, the more these matter.

Over to you: when you build an agent, which of these five takes the most work to get right?


How do you know if your AI app actually works?

You evaluate it. But most teams skip this step (or do it wrong) because “eval” feels vague. It’s not.

Every good eval is a 3-step recipe.

Step 1: Pick a task. AI systems have different capabilities and dimensions to evaluate. For LLMs, it can be safety or math capability, in RAGs it can be grounding and retrieval, Pick one.

Step 2: Collect eval data. For every task, gather inputs paired with the right answer or expected behavior. A safety set pairs risky prompts with “refuse.”

Step 3: Develop a grader. How do you decide if the output is good?

  • Use code-based graders (if/else, unit tests) for things with a clear correct answer and patch passing unit-tests.
  • Use model-based graders (LLM-as-judge) for subjective tasks like safety.
  • Use human graders for edge cases and anything where nuance matters more than throughput.

Most production evals combine all three. Code-based for what’s cheap to check. Model-based for scale. Human-based for what matters most.

Over to you: what’s the hardest thing about your task to grade, and which grader type do you use for it?


🚀 New Course: AI Evals in Practice Starts on Oct. 7

ByteByteGo has teamed up with Manjeet Singh, Senior Director at Salesforce, to bring you a live and hands-on course on building reliable evaluation systems for production AI agents.

Check it out Now

You’ll learn how to:

  • Design evals for quality, safety, reliability, cost, and latency
  • Build and validate LLM-as-a-Judge systems
  • Red-team agents for prompt injection and jailbreaks
  • Create meaningful eval datasets from real and synthetic data
  • Evaluate tool use, RAG, multi-step execution, and multi-agent handoffs
  • Run evals in CI/CD and production to catch regressions and drift
  • Turn failures into permanent regression tests

If you’re building AI agents and asking, “How do I know this is actually ready to ship?” — this course is for you.

Check it out Now

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论