---
title: "Autofill Node Processing"
description: "This document outlines the execution flow for the \"AutoFill -> ProcessNode\" process, which is initiated when a user requests to auto-fill a Statement of Applicability (SOA) document. The primary go..."
last_updated: "2026-05-06T07:30:58.159818+00:00"
canonical_url: "https://www.doc0.dev/docs/49c89830-117b-4def-8edc-b5bcc50766e0/technical/how-it-works/autofill-node-processing"
---

<details>
<summary>Relevant source files</summary>

The following files were used as context for generating this wiki page:

- [apps/api/src/soa/soa.controller.ts](https://github.com/blade47/comp/blob/main/apps/api/src/soa/soa.controller.ts)
- [apps/api/src/vector-store/lib/sync/sync-organization.ts](https://github.com/blade47/comp/blob/main/apps/api/src/vector-store/lib/sync/sync-organization.ts)
- [apps/api/src/vector-store/lib/sync/sync-policies.ts](https://github.com/blade47/comp/blob/main/apps/api/src/vector-store/lib/sync/sync-policies.ts)
- [apps/api/src/vector-store/lib/utils/extract-policy-text.ts](https://github.com/blade47/comp/blob/main/apps/api/src/vector-store/lib/utils/extract-policy-text.ts)
</details>

This document outlines the execution flow for the "AutoFill -> ProcessNode" process, which is initiated when a user requests to auto-fill a Statement of Applicability (SOA) document. The primary goal of this flow is to leverage the system's knowledge base to automatically generate answers for SOA questions.

The process begins with an API call to trigger the auto-fill operation. A critical initial step involves ensuring the vector database, which powers the AI's retrieval capabilities, is up-to-date with the latest organizational data, including policies. This synchronization ensures that the AI has access to the most relevant and current information when attempting to answer questions. The flow culminates in the extraction of plain text from rich policy content, making it suitable for embedding and subsequent retrieval by the AI.

### Step-by-step Narrative

The following steps detail the execution flow, from the initial API request to the deep-level text processing.

<Steps>
<Step>
### 1. Initiate Auto-Fill Request

The process starts with the `autoFill` method in the `SOAController`. This method is an HTTP POST endpoint designed to handle requests for automatically filling out an SOA document. It receives an `AutoFillSOADto` containing the document and organization IDs, along with the user's authentication context.

Upon invocation, the controller sets up Server-Sent Events (SSE) to stream real-time progress and answers back to the client. A crucial first action within this method is to trigger a synchronization of the organization's embeddings to ensure the vector database is current.

**Data Flow:**
*   **Input:** `AutoFillSOADto` (documentId, organizationId), `AuthContext` (userId), `Response` object for SSE.
*   **Output:** Initiates an SSE stream, potentially sending `progress`, `processing`, `answer`, `complete`, or `error` events.

<Callout variant="info" title="Error Handling">
The `autoFill` method includes a `try/catch` block around the call to `syncOrganizationEmbeddings`. If the synchronization fails, a warning is logged, but the auto-fill process attempts to continue, assuming some data might still be available in the vector DB. A broader `try/catch` wraps the entire `autoFill` logic, sending an SSE `error` event and logging any unhandled exceptions.
</Callout>

Sources: [apps/api/src/soa/soa.controller.ts:47-248](https://github.com/blade47/comp/blob/main/apps/api/src/soa/soa.controller.ts#L47-L248)
</Step>

<Step>
### 2. Synchronize Organization Embeddings

The `autoFill` method calls `syncOrganizationEmbeddings` to ensure the vector database is up-to-date. This function is responsible for orchestrating the synchronization of all relevant organizational data into the vector store. It implements a locking mechanism to prevent multiple concurrent sync operations for the same organization, ensuring data consistency and resource management.

If a sync for the given `organizationId` is already in progress, the function waits for the existing sync to complete rather than starting a new one.

**Data Flow:**
*   **Input:** `organizationId` (string) from the `AutoFillSOADto`.
*   **Output:** A Promise that resolves once the synchronization is complete, updating the vector store with the latest embeddings.

<Callout variant="info" title="Concurrency Control">
A `syncLocks` Map is used to manage ongoing synchronizations. If an entry exists for an `organizationId`, the function waits for the existing Promise to resolve, effectively serializing sync operations per organization.
</Callout>

Sources: [apps/api/src/vector-store/lib/sync/sync-organization.ts:21-51](https://github.com/blade47/comp/blob/main/apps/api/src/vector-store/lib/sync/sync-organization.ts#L21-L51)
</Step>

<Step>
### 3. Perform Core Synchronization Logic

The `syncOrganizationEmbeddings` function delegates the actual synchronization work to the internal `performSync` function. This function executes the core logic for updating the vector store with various types of organizational data.

`performSync` systematically fetches existing embeddings, then calls dedicated synchronization functions for policies, context entries, manual answers, and knowledge base documents. After processing these, it identifies and deletes any "orphaned" embeddings that no longer correspond to existing records in the database. Finally, it attempts to verify that newly created or updated embeddings are queryable, accounting for the eventual consistency of the vector store.

**Data Flow:**
*   **Input:** `organizationId` (string).
*   **Output:** Updates the vector store by upserting new/updated embeddings and deleting obsolete ones. Returns `Promise<void>`.

<Callout variant="warning" title="Eventual Consistency Handling">
The `verifyEmbeddingIsReady` function (called within `performSync`) uses a retry mechanism with exponential backoff. This is crucial for systems like Upstash Vector, which might have a delay between when an embedding is stored and when it becomes fully indexed and queryable.
</Callout>

Sources: [apps/api/src/vector-store/lib/sync/sync-organization.ts:54-135](https://github.com/blade47/comp/blob/main/apps/api/src/vector-store/lib/sync/sync-organization.ts#L54-L135)
</Step>

<Step>
### 4. Synchronize Policies

As part of `performSync`, the `syncPolicies` function is invoked to handle the synchronization of all published policies for the organization. It first fetches all relevant policies from the database using `fetchPolicies`.

The function then processes these policies in batches to optimize performance. For each policy in a batch, it calls `syncSinglePolicy` to manage the individual policy's embedding lifecycle. It aggregates statistics on created, updated, skipped, and failed policy synchronizations.

**Data Flow:**
*   **Input:** `organizationId` (string), `existingEmbeddingsMap` (Map of sourceId to `ExistingEmbedding[]`).
*   **Output:** `SyncStats` object (counts for created, updated, skipped, failed policies, and the ID of the last upserted embedding). Upserts policy embeddings into the vector store.

<Callout variant="info" title="Batch Processing">
Policies are processed in batches of `POLICY_BATCH_SIZE` (100) using `Promise.all` to allow for parallel execution, improving efficiency for organizations with many policies.
</Callout>

Sources: [apps/api/src/vector-store/lib/sync/sync-policies.ts:98-142](https://github.com/blade47/comp/blob/main/apps/api/src/vector-store/lib/sync/sync-policies.ts#L98-L142)
</Step>

<Step>
### 5. Synchronize a Single Policy

The `syncSinglePolicy` function is responsible for the detailed synchronization of an individual policy. It first checks if the policy's `updatedAt` timestamp has changed compared to existing embeddings. If not, it skips the policy.

If an update is needed, it deletes any old embeddings associated with this policy. It then calls `extractTextFromPolicy` to convert the policy's rich content into plain text. This plain text is then chunked into smaller, embeddable units, and these chunks are upserted into the vector store.

**Data Flow:**
*   **Input:** `policy` (PolicyData object), `existingEmbeddings` (array of `ExistingEmbedding` for this policy), `organizationId` (string).
*   **Output:** `SyncSingleResult` (status: 'created', 'updated', or 'skipped'; `lastEmbeddingId`). Modifies the vector store by deleting old embeddings and upserting new ones.

<Callout variant="success" title="Efficiency through Delta Sync">
The `needsUpdate` check prevents unnecessary re-embedding of policies that haven't changed, significantly reducing processing time and vector store operations.
</Callout>

Sources: [apps/api/src/vector-store/lib/sync/sync-policies.ts:63-95](https://github.com/blade47/comp/blob/main/apps/api/src/vector-store/lib/sync/sync-policies.ts#L63-L95)
</Step>

<Step>
### 6. Extract Text from Policy Content

Before a policy's content can be embedded, it must be converted into a clean, plain text format. The `extractTextFromPolicy` function takes a policy object, which typically contains rich text content in a structured format (like TipTap JSON), and transforms it into a single string of plain text.

It initializes the text with the policy's name and description, if available. Then, it iterates through the policy's content nodes, recursively calling `processNode` for each to extract text from the nested structure.

**Data Flow:**
*   **Input:** `policy` object (containing `name`, `description`, and `content` in TipTap JSON format).
*   **Output:** A single string representing the plain text content of the policy.

Sources: [apps/api/src/vector-store/lib/utils/extract-policy-text.ts:9-30](https://github.com/blade47/comp/blob/main/apps/api/src/vector-store/lib/utils/extract-policy-text.ts#L9-L30)
</Step>

<Step>
### 7. Process Individual Content Nodes

The `processNode` function is a recursive helper within `extractTextFromPolicy`. Its purpose is to traverse the tree-like structure of TipTap JSON content and extract plain text from various node types.

It handles different node types such as `text`, `heading`, `paragraph`, `bulletList`, and `orderedList`. For each type, it extracts the relevant text and formats it appropriately (e.g., adding bullet points for list items). If a node has child nodes, `processNode` calls itself recursively to process them, ensuring all nested content is captured.

**Data Flow:**
*   **Input:** A single TipTap `node` object (can be nested).
*   **Output:** A string representing the plain text content of that node and its children.

Sources: [apps/api/src/vector-store/lib/utils/extract-policy-text.ts:32-95](https://github.com/blade47/comp/blob/main/apps/api/src/vector-store/lib/utils/extract-policy-text.ts#L32-L95)
</Step>
</Steps>

### Sequence Diagram



### Flowchart



### Key Observations

*   **Cross-Module Boundaries:** This flow demonstrates significant interaction across different modules:
    *   The `SOAController` (API layer) initiates the process.
    *   The `vector-store/lib/sync` module handles the core synchronization logic with the vector database.
    *   The `vector-store/lib/utils` module provides utility functions for text extraction.
    *   Interactions with the `Database` (via Prisma ORM) and the external `VectorStore` (e.g., Upstash Vector) are central to the data flow.

*   **Potential Failure Points:**
    *   **Authentication:** The `autoFill` endpoint requires user authentication, failing early if not met.
    *   **Vector Store Sync:** `syncOrganizationEmbeddings` is a critical dependency. While `autoFill` attempts to proceed if sync fails (logging a warning), this could lead to less accurate auto-fill results if the vector store is outdated.
    *   **Database Connectivity:** Failures to fetch policies, documents, or save answers will halt the process.
    *   **Vector Store Operations:** Issues during upserting, querying, or deleting embeddings can cause sync failures.
    *   **Content Extraction:** Malformed or unexpected policy content (TipTap JSON) could lead to incomplete or incorrect text extraction, impacting embedding quality.
    *   **Concurrency:** The `syncLocks` mechanism in `syncOrganizationEmbeddings` is vital to prevent race conditions and ensure data integrity during concurrent sync requests for the same organization.

*   **Performance Considerations:**
    *   **Initial Sync Latency:** The `syncOrganizationEmbeddings` call at the beginning of `autoFill` can introduce noticeable latency, especially for organizations with a large volume of data (policies, documents) that need to be processed. This is a trade-off for ensuring the AI has the freshest data.
    *   **Batch Processing:** `syncPolicies` uses batch processing (`POLICY_BATCH_SIZE`) and `Promise.all` to parallelize policy synchronization, improving efficiency.
    *   **Delta Sync:** The `needsUpdate` check in `syncSinglePolicy` is a key optimization, preventing unnecessary re-embedding of unchanged policies.
    *   **Eventual Consistency:** The `verifyEmbeddingIsReady` function, with its exponential backoff retries, adds necessary delays to ensure data is queryable but can extend the overall sync time.
    *   **SSE for Responsiveness:** Using Server-Sent Events (SSE) for `autoFill` provides a better user experience by streaming progress and answers, masking some of the backend processing time.

## Sitemap

See the full [sitemap](https://www.doc0.dev/docs/49c89830-117b-4def-8edc-b5bcc50766e0/llms.txt) for all pages in this wiki.
