Clearwater, Florida · Hyderabad, India

Building RAG on Amazon Bedrock Knowledge Bases

How to build retrieval-augmented generation with Knowledge Bases for Amazon Bedrock — data sources, chunking, vector stores, metadata filtering, evaluation and security.

Retrieval-augmented generation (RAG) is how you get a foundation model to answer questions from your own documents — accurately and with citations. Knowledge Bases for Amazon Bedrock turns much of the RAG pipeline into a managed service. Here is how to design one that works well in production.

How Knowledge Bases work

You point a knowledge base at your content, choose an embedding model and a vector store, and Bedrock handles ingestion: parsing documents, splitting them into chunks, creating embeddings and storing them. At query time, the user’s question is embedded, the most relevant chunks are retrieved and passed to a model, and the answer is returned with references to the source passages.

Step 1: Prepare your content

Quality in, quality out. Start with a well-defined, owned set of documents rather than “everything on the file share”. Remove duplicates and outdated versions, and make sure each document has an owner who will keep it current. Content commonly lives in Amazon S3, and connectors can bring in sources such as SharePoint, Confluence, Salesforce and websites.

Step 2: Choose chunking and parsing

How documents are split has a large effect on answer quality. Fixed-size chunks are simple; hierarchical or semantic chunking often works better for long, structured documents such as policies and manuals. For PDFs with tables and complex layouts, use advanced parsing so that tables and headings survive ingestion intact.

Step 3: Pick a vector store

Knowledge Bases supports several vector stores, including Amazon OpenSearch Serverless and Amazon Aurora PostgreSQL with pgvector, among others. Choose based on scale, existing skills and cost. For most first deployments, the managed default is the quickest path; teams already running Aurora often prefer pgvector.

Step 4: Use metadata and filtering

Attach metadata — department, product, region, document type, effective date — to your files. At query time you can filter on it, so a user in one region only receives that region’s policy, or answers come only from current documents. Metadata filtering is also a key building block for enforcing access rules.

Step 5: Secure it

  • Respect document permissions: separate knowledge bases or metadata filters per audience, tied to the user’s identity.
  • Use Guardrails to block off-topic requests and redact sensitive information.
  • Keep traffic private with VPC endpoints, encrypt with KMS and log access with CloudTrail.

Step 6: Evaluate before and after launch

Create a test set of 50–200 real questions with expected answers and sources. Measure whether the right passages are retrieved, whether answers are correct and complete, and whether citations support them. Re-run the evaluation whenever you change chunking, models, prompts or content — and review real user feedback weekly in the early months.

Common pitfalls

  • Indexing outdated or conflicting documents, so the model gets contradictory context.
  • Chunks that are too small to carry meaning, or too large to be precise.
  • No instruction to say “I don’t know” when the answer isn’t in the retrieved context.
  • Skipping evaluation and relying on a handful of demo questions.

Need help? Our Amazon Bedrock specialists build and tune knowledge bases for production use.

Frequently asked questions

What vector stores do Bedrock Knowledge Bases support?

Supported options include Amazon OpenSearch Serverless and Amazon Aurora PostgreSQL with pgvector, among others. Check AWS documentation for the current list.

Can Bedrock Knowledge Bases cite sources?

Yes. Responses can include references to the retrieved source passages, which helps users verify answers.

How do I restrict which documents a user can see?

Common approaches are separate knowledge bases per audience or metadata filtering applied at query time based on the user's identity and permissions.

Ready to build what’s next?

Tell us about your project. Our consultants in Florida and Hyderabad will get back to you within one business day.

Start a conversation