# AI Knowledge Base

Give an agent something true to answer from. You point a knowledge base at a URL; Stark Infra crawls the site, converts every page to Markdown and indexes it so the relevant passages can be pulled back in milliseconds.

This is what separates an assistant that helps from one that improvises. An agent with a knowledge base retrieves the passages that match the question and answers from them, instead of guessing from what the model happened to memorize.

Retrieval is heading-level, not file-level: a question matches the section that answers it, so the agent reads a few hundred relevant words instead of your whole site. Watch the crawl fill up page by page while it runs, and the base keeps itself current from then on.

NOTE: Read [Core Concepts](https://docs.starkinfra.com/get-started/core-concepts.md) before continuing this guide.

**RESOURCE SUMMARY**

**AI Knowledge Base**A crawled and indexed website your agents answer from.

## Available Languages

Code samples on this page are in Python. The same content is available with samples in:

- [Python](https://docs.starkinfra.com/get-started/ai-knowledge-base-python.md)
- [Node.js](https://docs.starkinfra.com/get-started/ai-knowledge-base-node.md)
- [PHP](https://docs.starkinfra.com/get-started/ai-knowledge-base-php.md)
- [Java](https://docs.starkinfra.com/get-started/ai-knowledge-base-java.md)
- [Ruby](https://docs.starkinfra.com/get-started/ai-knowledge-base-ruby.md)
- [Elixir](https://docs.starkinfra.com/get-started/ai-knowledge-base-elixir.md)
- [.NET](https://docs.starkinfra.com/get-started/ai-knowledge-base-dotnet.md)
- [Go](https://docs.starkinfra.com/get-started/ai-knowledge-base-go.md)
- [Clojure](https://docs.starkinfra.com/get-started/ai-knowledge-base-clojure.md)
- [cURL](https://docs.starkinfra.com/get-started/ai-knowledge-base-curl.md)

## Setup

For each environment (Sandbox or Production):

1. Create a workspace at Stark Infra and generate your ECDSA keys.

2. Get in touch with your account manager to enable the AI products on your workspace.

3. Have the URL you want indexed ready, and make sure it is reachable from the public internet — the crawler reads your site the same way a browser does.

## Typical flow

**1.** Create a knowledge base with the `rootUrl` you want indexed. The call returns immediately with `status=processing`.

**2.** Poll `GET /v2/ai-knowledge-base/:id` until the status leaves `processing`. Small sites finish in minutes; large ones take longer.

**3.** Optionally watch the crawl page by page with the hosts endpoint, which shows every URL seen and whether it was indexed.

**4.** Once the base reaches `success`, attach it to an [AI Agent](https://docs.starkinfra.com/get-started/ai-agent.md) through `knowledgeBaseIds`. From then on the agent retrieves from it on every message.

**5.** Nothing else to schedule: the base is re-crawled daily, so the index tracks your site as it changes.

## What gets indexed

The crawler seeds itself from three places: the root page you gave it, the site's `sitemap.xml`, and its `llms.txt` when one exists. Publishing either of those last two is the cheapest way to control exactly what ends up in the index.

From those seeds it follows links. With `isRecursive` at its default `true`, it may cross into other subdomains of the same registered domain — `docs.example.com` can reach `api.example.com`. Send `false` to keep it on the root host.

Pages are indexed at heading level, not file level. A question matches the section that answers it, so the agent reads a few hundred relevant words instead of a few hundred kilobytes of site.

## Use cases

**Support that cites your policy:** index your help center so the agent answers refunds and shipping from the page your team actually maintains.

**Developer assistant:** index your API documentation and let the agent answer integration questions with your real endpoints and parameters.

**Multilingual support:** index the same documentation once and let agents answer from it in whichever language the customer writes in.

**Onboarding assistant:** index your integration guides so a new customer gets answers with your real endpoints and parameters.

## AI Knowledge Base Overview

Here we show you how to index a site, follow the crawl while it runs, and read the passages back — either through an agent or directly.

### Creating a knowledge base

`POST /v2/ai-knowledge-base`

Give it a name you will recognize and the `rootUrl` to start from. The call returns immediately with `status=processing` — crawling happens in the background.

Use `tags` to keep bases organized once you have several: one per product, one per language, one per audience.

**Request**

```python
import starkinfra

knowledge_base = starkinfra.aiknowledgebase.create(
    starkinfra.AiKnowledgeBase(
        name="Product Documentation",
        root_url="https://docs.starkinfra.com",
        tags=["support", "public"]
    )
)

print(knowledge_base)
```

**Response**

```python
AiKnowledgeBase(
    created=2022-01-01 00:00:00,
    id=6767676767676767,
    is_recursive=True,
    name=Product Documentation,
    root_url=https://docs.starkinfra.com,
    status=processing,
    tags=['support', 'public'],
    updated=2022-01-01 00:00:00
)
```

### Waiting for the crawl

`GET /v2/ai-knowledge-base/:id`

Poll the base until `status` leaves `processing`. There is no webhook for this transition.

`success` means at least one page was indexed and the base is ready to answer. `failed` means none was — almost always because the root URL was unreachable, or because the site exposed no links the crawler could follow.

Attaching a base that is still processing to an agent is allowed. The agent simply retrieves nothing from it until the index is ready.

**Request**

```python
import starkinfra

knowledge_base = starkinfra.aiknowledgebase.get("6767676767676767")

print(knowledge_base)
```

**Response**

```python
AiKnowledgeBase(
    created=2022-01-01 00:00:00,
    id=6767676767676767,
    is_recursive=True,
    name=Product Documentation,
    root_url=https://docs.starkinfra.com,
    status=success,
    tags=['support', 'public'],
    updated=2022-01-01 00:10:00
)
```

### Watching the crawl page by page

`GET /v2/ai-knowledge-base/:id/hosts`

Get every URL the crawler has seen, grouped by host, with the Markdown file stored for each one and its status: `pending` while still queued, `success` once stored, `failed` when it could not be read.

This is the call to reach for when a base finishes with less content than you expected: it names the pages that failed, so you can tell a blocked crawler from a site that simply has fewer pages than you thought.

**Request**

```python
import starkinfra

hosts = starkinfra.aiknowledgebase.hosts("6767676767676767")

for host, pages in hosts.items():
    print(host)
    for page in pages:
        print(page)
```

**Response**

```python
docs.starkinfra.com
{'original_url': 'https://docs.starkinfra.com/get-started', 'status': 'success', 'storage_url': 'https://storage.googleapis.com/ai-knowledge/6767676767676767/get-started.md'}
{'original_url': 'https://docs.starkinfra.com/api', 'status': 'success', 'storage_url': 'https://storage.googleapis.com/ai-knowledge/6767676767676767/api.md'}
{'original_url': 'https://docs.starkinfra.com/sandbox', 'status': 'pending', 'storage_url': None}
```
