GuideUpdated 2026-09-22

Kimi K3 Lands on Amazon Bedrock With Prompt Caching

A large context window and discounted cache hits can help repeated knowledge work, but buyers still need task-level quality, residency, and cost tests.

By DiscoverAI Editorial TeamReviewed by DiscoverAI Editorial Review2 min readContent & SearchHow we evaluate
Paper-cut editorial illustration of a multimodal model reading code documents and images through a million-token corridor with reusable prompt blocks and cost meters
Original DiscoverAI editorial illustration. Editorial illustration: context capacity and cache discounts create value only when representative tasks remain accurate and reviewable.

Bottom line

AWS added Moonshot AI's Kimi K3 to Amazon Bedrock with native vision, a one-million-token context window, OpenAI-compatible APIs, and explicit prompt caching.

Editorial accountability

Who checked this guide

Meet the editorial team →
Evaluation type
Research-based verification
Last materially checked
Evidence
4 listed sources

Hands-on testing is identified explicitly. Research-based coverage uses cited product documentation and other named sources; it does not imply every paid plan was used. Read the full methodology.

Editorial basis

What this guidance is based on

Editorial basis
Source-led analysis
Primary references
4
Products covered
2
Last checked
2026-09-22

Important limits

  • Model capability and efficiency claims originate from AWS and Moonshot AI.
  • Task quality, caching economics, availability, and compliance fit require buyer testing.
In this guide
  1. Short answer
  2. What is new
  3. Why prompt caching matters
  4. What the launch does not prove
  5. A fair migration test

Short answer

AWS made Moonshot AI's Kimi K3 available on Amazon Bedrock on September 18, 2026, positioning it for coding and knowledge work with native vision, a one-million-token context window, and explicit prompt caching. AWS says Bedrock requests stay inside its data boundary, are not shared with the model provider, and use zero data retention for inference. Those platform claims do not establish accuracy, software quality, or fitness for a regulated workload.

What is new

Kimi K3 can be called through Bedrock's native APIs or OpenAI-compatible Responses and Chat Completions interfaces. Global and U.S. geographic inference profiles are listed. The explicit cache lets developers mark a reusable prompt prefix of at least 1,024 tokens; writes cost more, while matching hits receive discounted input pricing and do not count against input-token-per-minute quotas during the documented cache period.

Why prompt caching matters

Coding assistants and research agents repeatedly resend repository guidance, tool definitions, schemas, and reference documents. A cache can reduce repeated-input latency and cost when the prefix remains identical and requests arrive within the retention window. It adds little when context changes constantly, requests are sparse, or cache misses dominate.

What the launch does not prove

Parameter count and context length are capacity descriptions, not outcome measures. Long inputs can still produce missed evidence, incorrect code, security flaws, or expensive outputs. Moonshot and AWS performance claims need validation on the buyer's own languages, repositories, images, tools, and failure cases.

A fair migration test

Use 50 representative tasks and compare Kimi K3 with the current model under identical retrieval, tools, prompts, review, and regional constraints. Record accepted-task rate, citations, test pass rate, security findings, latency, cache hit rate, input/output cost, reviewer time, and rollback behavior.

Sources and verification

Product details and claims were checked against the following primary sources.

Frequently asked questions

What is Kimi K3?

Kimi K3 is Moonshot AI's open-weight multimodal model, offered on Amazon Bedrock for coding and knowledge workflows.

How large is Kimi K3's context window?

AWS lists a one-million-token context window, but usable retrieval and reasoning quality must be tested on representative long inputs.

How does Kimi K3 prompt caching work on Bedrock?

Developers mark a stable prefix of at least 1,024 tokens; cache writes cost more, while matching hits can lower input cost and latency for the documented cache duration.

Does Bedrock send prompts to Moonshot AI?

AWS states that open-weight model inference data remains within the AWS data boundary and is not shared with the model provider; buyers should verify the applicable contract and configuration.

Found this useful?

Get the next one in your inbox.

One five-minute briefing a week: a meaningful change, a practical workflow, and a clearer tool decision—already filtered for lean teams.

Free · one email a week · unsubscribe any time

Read next

More on Content & Search