System Architecture
Core Features
Data Management
Frontend Components
Extensibility
The following files were used as context for generating this wiki page:
This document outlines the execution flow for the "AutoFill -> ProcessNode" process, which is initiated when a user requests to auto-fill a Statement of Applicability (SOA) document. The primary goal of this flow is to leverage the system's knowledge base to automatically generate answers for SOA questions.
The process begins with an API call to trigger the auto-fill operation. A critical initial step involves ensuring the vector database, which powers the AI's retrieval capabilities, is up-to-date with the latest organizational data, including policies. This synchronization ensures that the AI has access to the most relevant and current information when attempting to answer questions. The flow culminates in the extraction of plain text from rich policy content, making it suitable for embedding and subsequent retrieval by the AI.
The following steps detail the execution flow, from the initial API request to the deep-level text processing.
The process starts with the autoFill method in the . This method is an HTTP POST endpoint designed to handle requests for automatically filling out an SOA document. It receives an containing the document and organization IDs, along with the user's authentication context.
SOAControllerAutoFillSOADtoUpon invocation, the controller sets up Server-Sent Events (SSE) to stream real-time progress and answers back to the client. A crucial first action within this method is to trigger a synchronization of the organization's embeddings to ensure the vector database is current.
Data Flow:
AutoFillSOADto (documentId, organizationId), AuthContext (userId), Response object for SSE.progress, processing, answer, complete, or error events.The autoFill method includes a try/catch block around the call to syncOrganizationEmbeddings. If the synchronization fails, a warning is logged, but the auto-fill process attempts to continue, assuming some data might still be available in the vector DB. A broader try/catch wraps the entire autoFill logic, sending an SSE error event and logging any unhandled exceptions.
The autoFill method calls syncOrganizationEmbeddings to ensure the vector database is up-to-date. This function is responsible for orchestrating the synchronization of all relevant organizational data into the vector store. It implements a locking mechanism to prevent multiple concurrent sync operations for the same organization, ensuring data consistency and resource management.
If a sync for the given organizationId is already in progress, the function waits for the existing sync to complete rather than starting a new one.
Data Flow:
organizationId (string) from the AutoFillSOADto.A syncLocks Map is used to manage ongoing synchronizations. If an entry exists for an organizationId, the function waits for the existing Promise to resolve, effectively serializing sync operations per organization.
Sources: apps/api/src/vector-store/lib/sync/sync-organization.ts:21-51
The syncOrganizationEmbeddings function delegates the actual synchronization work to the internal performSync function. This function executes the core logic for updating the vector store with various types of organizational data.
performSync systematically fetches existing embeddings, then calls dedicated synchronization functions for policies, context entries, manual answers, and knowledge base documents. After processing these, it identifies and deletes any "orphaned" embeddings that no longer correspond to existing records in the database. Finally, it attempts to verify that newly created or updated embeddings are queryable, accounting for the eventual consistency of the vector store.
Data Flow:
organizationId (string).Promise<void>.The verifyEmbeddingIsReady function (called within performSync) uses a retry mechanism with exponential backoff. This is crucial for systems like Upstash Vector, which might have a delay between when an embedding is stored and when it becomes fully indexed and queryable.
Sources: apps/api/src/vector-store/lib/sync/sync-organization.ts:54-135
As part of performSync, the syncPolicies function is invoked to handle the synchronization of all published policies for the organization. It first fetches all relevant policies from the database using fetchPolicies.
The function then processes these policies in batches to optimize performance. For each policy in a batch, it calls syncSinglePolicy to manage the individual policy's embedding lifecycle. It aggregates statistics on created, updated, skipped, and failed policy synchronizations.
Data Flow:
organizationId (string), existingEmbeddingsMap (Map of sourceId to ExistingEmbedding[]).SyncStats object (counts for created, updated, skipped, failed policies, and the ID of the last upserted embedding). Upserts policy embeddings into the vector store.Policies are processed in batches of POLICY_BATCH_SIZE (100) using Promise.all to allow for parallel execution, improving efficiency for organizations with many policies.
Sources: apps/api/src/vector-store/lib/sync/sync-policies.ts:98-142
The syncSinglePolicy function is responsible for the detailed synchronization of an individual policy. It first checks if the policy's updatedAt timestamp has changed compared to existing embeddings. If not, it skips the policy.
If an update is needed, it deletes any old embeddings associated with this policy. It then calls extractTextFromPolicy to convert the policy's rich content into plain text. This plain text is then chunked into smaller, embeddable units, and these chunks are upserted into the vector store.
Data Flow:
policy (PolicyData object), existingEmbeddings (array of ExistingEmbedding for this policy), organizationId (string).SyncSingleResult (status: 'created', 'updated', or 'skipped'; lastEmbeddingId). Modifies the vector store by deleting old embeddings and upserting new ones.The needsUpdate check prevents unnecessary re-embedding of policies that haven't changed, significantly reducing processing time and vector store operations.
Sources: apps/api/src/vector-store/lib/sync/sync-policies.ts:63-95
Before a policy's content can be embedded, it must be converted into a clean, plain text format. The extractTextFromPolicy function takes a policy object, which typically contains rich text content in a structured format (like TipTap JSON), and transforms it into a single string of plain text.
It initializes the text with the policy's name and description, if available. Then, it iterates through the policy's content nodes, recursively calling processNode for each to extract text from the nested structure.
Data Flow:
policy object (containing name, description, and content in TipTap JSON format).Sources: apps/api/src/vector-store/lib/utils/extract-policy-text.ts:9-30
The processNode function is a recursive helper within extractTextFromPolicy. Its purpose is to traverse the tree-like structure of TipTap JSON content and extract plain text from various node types.
It handles different node types such as text, heading, paragraph, bulletList, and orderedList. For each type, it extracts the relevant text and formats it appropriately (e.g., adding bullet points for list items). If a node has child nodes, processNode calls itself recursively to process them, ensuring all nested content is captured.
Data Flow:
node object (can be nested).Sources: apps/api/src/vector-store/lib/utils/extract-policy-text.ts:32-95
Cross-Module Boundaries: This flow demonstrates significant interaction across different modules:
SOAController (API layer) initiates the process.vector-store/lib/sync module handles the core synchronization logic with the vector database.vector-store/lib/utils module provides utility functions for text extraction.Database (via Prisma ORM) and the external VectorStore (e.g., Upstash Vector) are central to the data flow.Potential Failure Points:
autoFill endpoint requires user authentication, failing early if not met.syncOrganizationEmbeddings is a critical dependency. While autoFill attempts to proceed if sync fails (logging a warning), this could lead to less accurate auto-fill results if the vector store is outdated.syncLocks mechanism in syncOrganizationEmbeddings is vital to prevent race conditions and ensure data integrity during concurrent sync requests for the same organization.Performance Considerations:
syncOrganizationEmbeddings call at the beginning of autoFill can introduce noticeable latency, especially for organizations with a large volume of data (policies, documents) that need to be processed. This is a trade-off for ensuring the AI has the freshest data.syncPolicies uses batch processing (POLICY_BATCH_SIZE) and Promise.all to parallelize policy synchronization, improving efficiency.needsUpdate check in syncSinglePolicy is a key optimization, preventing unnecessary re-embedding of unchanged policies.verifyEmbeddingIsReady function, with its exponential backoff retries, adds necessary delays to ensure data is queryable but can extend the overall sync time.autoFill provides a better user experience by streaming progress and answers, masking some of the backend processing time.