Enterprise AI is often presented as a model-selection problem. Pick a capable model, connect a few sources, add a chat interface, and the product is done.
That architecture works until the software has to become part of the operating environment of a real enterprise.
Then the questions change. Where does sensitive data execute? Which identity system remains authoritative? Can a customer run the platform inside its own cloud boundary? What happens when a model endpoint is unavailable, a region degrades, or an integration returns inconsistent data? Can an agent act without becoming a second authorization system? Can the platform evolve without forcing customers to surrender control of infrastructure, models, networking, or data residency?
Those questions shaped Facthory much more than the question of which LLM happens to be strongest this month.
Facthory is being built as an operational intelligence platform: a distributed system that connects enterprise knowledge, systems, people, multimodal data, and governed AI agents into one shared operating context. The model is a component of that system. It is not the system.
This post describes the architecture we are building toward. It focuses on public system boundaries, deployment models, reliability principles, and security guarantees. It intentionally does not describe proprietary context construction, retrieval ranking, index design, scheduling, caching, model-routing heuristics, optimization loops, or other performance-critical mechanics.
A useful way to understand Facthory is to separate the platform into two architectural responsibilities: a control plane and one or more data planes.
The control plane coordinates the platform. It owns tenant-level configuration, identity integration, policy, lifecycle, deployment state, governance metadata, and the operational view required to manage the system.
The data plane is where customer-sensitive execution belongs. It connects to enterprise systems, processes multimodal information, executes governed workloads, talks to approved model endpoints, and stores customer-controlled operational data according to the deployment boundary selected by the enterprise.
This distinction is fundamental. It lets us centralize governance without requiring every enterprise to centralize its data inside our infrastructure.

The important part is not the boxes. The important part is what does not cross the boundary by default.
A customer using Facthory in a customer-owned deployment should not have to move bulk operational data into a vendor SaaS environment merely to receive platform intelligence. The control plane can coordinate policy, lifecycle, and system state while sensitive execution remains inside the customer environment.
That is the architectural foundation for sovereignty.
Enterprises do not share one risk model. A mid-market manufacturer, a defense supplier, a pharmaceutical company, and a public-sector organization may all want the same product capabilities while requiring very different infrastructure boundaries.
Facthory is therefore designed around several deployment envelopes rather than one fixed hosting assumption.

| Deployment model | Primary control boundary | Designed for |
|---|---|---|
| Managed Enterprise | Dedicated Facthory-operated environment | Organizations that want enterprise isolation without operating the infrastructure themselves |
| BYOC | Customer cloud subscription and network | Enterprises that want Facthory inside their own Azure or cloud governance boundary |
| Private Cloud | Dedicated infrastructure and private connectivity | Organizations with stricter isolation, network, or compliance requirements |
| On-Premises | Customer-controlled local infrastructure | Restricted environments where data and execution must remain inside the local estate |
BYOC is particularly important because data residency alone is not sovereignty. An enterprise may also need authority over networking, keys, cloud policy, model endpoints, cost allocation, release windows, and infrastructure operations.
Putting the data plane inside the customer's environment changes the relationship. Facthory becomes software operating inside an enterprise security boundary rather than a remote assistant that the enterprise must continuously export context to.
Facthory processes long-running, multimodal, and operational workloads. A video may take time to process. A document pipeline may involve multiple stages. An enterprise investigation may continue for hours or days. A model endpoint may slow down. A connector may temporarily fail. A factory network may be unreliable.
Trying to represent all of that as a synchronous request chain creates fragile software.
Our architectural direction is therefore event-driven and loosely coupled. Services are aligned to durable business capabilities, own their lifecycle and state boundaries, and communicate asynchronously where the workload does not require immediate request-response semantics.
NATS is part of the event fabric in the current platform direction. PostgreSQL provides durable transactional state where relational guarantees matter. Redis is used for ephemeral and coordination-oriented state where appropriate. Large multimodal objects belong in object storage rather than being pushed through application processes. Kubernetes provides the common scheduling and isolation substrate across deployment models.
None of those technologies is the moat.
The moat is the behavior of the system built around them: how work is partitioned, governed, recovered, observed, upgraded, and connected to enterprise context without collapsing into a tightly coupled chain.
A reliable system is not one in which nothing fails. It is one in which failures remain bounded.
Facthory's engineering principles include:
idempotent processing where operations may be retried
exponential backoff rather than retry storms
circuit breakers around unstable dependencies
durable asynchronous work for long-running operations
dead-letter handling and recoverable failure paths
graceful degradation when a non-critical capability is unavailable
explicit health, telemetry, and audit signals
rollback-first release engineering
The objective is to stop one degraded dependency from turning into a platform-wide outage.
Kubernetes can restart a process. That is useful, but process restart is the smallest reliability problem an enterprise platform has to solve.
The larger questions are about failure domains: a node, an availability zone, a managed database instance, a network path, an external model provider, or an entire region.
The target Facthory topology treats stateless application capacity as replaceable and spreads it across zones. Stateful components use replication and recovery semantics appropriate to their role. Regional deployments are designed so that critical platform responsibilities can be recovered or failed over without assuming that one machine, one pod, or one zone is permanent.

This diagram is intentionally abstract because replication semantics differ by component and deployment model. The design principle is consistent: availability is created by removing single failure domains, not by making one server bigger.
For customers, this matters because the platform is expected to become part of daily operations. For investors, it matters because enterprise software value is created by everything required to operate reliably after the demo works.
A login screen is not an enterprise security architecture.
Facthory treats identity, authorization, network boundaries, workload identity, secrets, model access, and auditability as separate controls that reinforce each other.
The enterprise identity provider remains authoritative for user identity. Deployments can integrate SSO and MFA. Workloads use managed or short-lived identities where the environment supports them. Secrets are stored outside application code. Transport is encrypted. Data stores and object storage are encrypted. Private connectivity can be used for enterprise systems and approved model endpoints.
Most importantly, the model is never the authority.
A model may propose an action. It does not grant itself permission to perform that action.

This separation is also why Facthory's agent architecture can become more capable without becoming less governable. Agent instructions, learned procedures, model outputs, and tool permissions are different concerns. Policy remains independently authoritative.
Depending on deployment mode, the broader security envelope can include:
enterprise SSO and MFA
role-based and attribute-aware authorization
tenant and source-level access boundaries
TLS for service and user traffic
private endpoints and restricted ingress paths
managed workload identity and vault-backed secrets
encryption for databases and object storage
data-retention and deletion controls
evidence and provenance records
agent activity histories and human approval gates
model governance and approved endpoint policies
customer-controlled regional placement
This is the difference between adding AI to an application and treating AI as governed enterprise infrastructure.
When enterprises say they need sovereign AI, they often mean more than where a file is stored.
We think about sovereignty across at least three dimensions.
Where data is stored, processed, retained, and deleted. In BYOC or private deployments, the customer can keep the sensitive data plane inside its own cloud and network boundary.
Which models are allowed, where inference runs, which providers are approved, and whether customer-controlled or private model endpoints are required. Facthory is designed around a model-governance layer rather than a hard dependency on one model vendor.
Where workloads execute, which network paths they may use, which enterprise systems they can reach, and which infrastructure policies govern them.
These dimensions matter because an enterprise can satisfy data residency while still losing control over inference, networking, or operational execution. Facthory's deployment model is intended to keep those decisions explicit.
Operational knowledge is not a collection of text documents.
It lives in video, images, audio, maintenance records, procedures, system events, quality data, dashboards, technical files, conversations, and the judgment of people who understand how the operation actually behaves.
That creates data-gravity problems that generic document assistants rarely have to solve at the same depth.
Our architectural principles are straightforward:
Large objects belong in durable object storage rather than passing repeatedly through application services.
Processing should happen close to the data boundary whenever possible.
Metadata, provenance, ownership, validation state, and lineage are first-class product data.
Transactional state belongs in durable stores such as PostgreSQL.
Ephemeral acceleration and coordination state can use systems such as Redis, but should not become an accidental source of truth.
Long-running processing should remain recoverable when a browser closes or a worker restarts.
The proprietary part is what happens inside the intelligence layer after these foundations are in place. We deliberately do not publish how Facthory constructs context, ranks evidence, combines modalities, schedules expensive processing, or optimizes retrieval and inference paths.
Those mechanisms are product IP, not architecture documentation.
A system is not production-grade because its source code looks clean. It becomes production-grade when the delivery process continuously proves that changes preserve contracts, behavior, architecture, and recoverability.
Facthory's engineering model uses multiple quality layers rather than one test suite.
Current platform engineering includes contract and schema validation, service-level unit testing, integration testing against real infrastructure such as NATS and PostgreSQL, behavior-driven tests, architecture fitness checks, mutation testing, and failure-injection or chaos exercises for selected paths.
Expensive validation belongs where it can run without turning every pull request into a queue. Fast structural and behavioral checks belong in the change path. Broader mutation and chaos runs can continuously measure system health outside the developer critical path.

GitOps is important here because infrastructure and application state should be reviewable and reproducible. The target is a platform where environments are created from declared state, changes are visible in version control, and progressive rollout can stop a bad release before it becomes a fleet-wide event.
This is less glamorous than an agent demo. It is also the work that separates an enterprise platform from a prototype.
Facthory deliberately uses widely understood infrastructure where it gives us leverage. Reinventing a database, message broker, or container scheduler would not make the product more differentiated.
| Foundation | Role in the platform | Why it is not the moat |
|---|---|---|
| Kubernetes | Portable scheduling, scaling and workload isolation across managed and customer-controlled environments | The differentiation is the platform behavior and deployment contract above the scheduler |
| PostgreSQL | Durable relational state where transactional guarantees matter | The schema alone does not create operational intelligence |
| Redis | Ephemeral state, acceleration and coordination where appropriate | It is an implementation primitive, not the source of product intelligence |
| NATS | Cloud-native event transport for asynchronous, recoverable workflows | The valuable part is how domain work is partitioned, governed and recovered |
| Object storage | Durable storage for large multimodal artifacts and derived data | Value comes from how evidence, provenance and operational context are built around the objects |
| API gateway and private ingress | Controlled north-south access, routing and policy enforcement points | The gateway does not define the product security model |
| GitOps | Reproducible environments and reviewable desired state | The deployment system is only useful because the platform is designed to be portable and declarative |
This is intentional. We want the commodity layers to be replaceable. We want the difficult product behavior to remain ours.
A general-purpose SaaS copilot and an operational intelligence platform can both contain a chat interface. That similarity ends quickly once the architecture is examined.
| Design question | Typical general-purpose SaaS copilot design center | Facthory design center |
|---|---|---|
| Primary abstraction | Assistant session | Persistent enterprise operational context |
| Deployment boundary | Primarily vendor-operated SaaS | Managed, BYOC, private cloud, or on-premises |
| Data control | Optimized around the provider ecosystem | Customer-controlled data plane where required |
| Model strategy | Usually centered on the provider's model ecosystem | Governed model choice and customer-controlled endpoints |
| Work duration | Conversation-oriented assistance | Durable work that may continue across teams, systems, and long-running agent activity |
| Authority | Application permission model | Independent identity, policy, capability, approval, and tool-execution boundaries |
| Enterprise context | Connected content and application context | Operational model spanning people, processes, systems, evidence, assets, decisions, and outcomes |
| Failure model | Application availability | Distributed failure domains across services, models, connectors, zones, regions, and customer environments |
This is not an argument that every enterprise needs the maximum architecture on day one. It is an argument that the architecture must have somewhere to grow when the first successful AI use case becomes infrastructure for the rest of the company.
The easiest architecture to build is the one that assumes the product will remain small.
One tenant. One model provider. One region. One database. One synchronous request path. One set of administrators. One cloud account. One security boundary.
That design is attractive because the first demo arrives quickly.
It also creates a ceiling.
Facthory is being designed for the opposite scenario: the first use case succeeds, more business units arrive, more agents run concurrently, more systems are connected, customers demand private deployment, security teams require stronger controls, data volumes increase, and the platform becomes important enough that downtime or an uncontrolled action has a real business cost.
That is why the architecture includes service autonomy, event-driven work, control-plane and data-plane separation, enterprise identity, policy boundaries, customer-owned deployment, zonal resilience, regional recovery, GitOps, progressive delivery, auditability, and failure testing.
It is a lot of engineering.
That is the point.
An architecture blog should explain the shape of the system without publishing the machinery that creates product advantage.
We therefore do not describe the internal algorithms and implementation details behind:
multimodal context construction
retrieval and evidence ranking
graph traversal and relationship scoring
processing and workload scheduling
cache hierarchy and latency optimization
model selection and routing heuristics
evaluation and fallback logic
context compression and token economics
self-improvement mechanisms beyond their governance boundaries
internal data layouts and index structures
Those areas evolve quickly and represent a meaningful part of the engineering investment in Facthory.
The public architecture is the contract. The private machinery is how we make that contract useful at enterprise scale.
The architectural thesis behind Facthory is simple.
Enterprise AI should not require an organization to choose between capability and control.
A customer should be able to run advanced multimodal and agentic workloads while retaining authority over identity, data, models, networking, regions, infrastructure, and consequential actions. The platform should survive dependency failures, scale beyond a single application instance, recover from bad releases, preserve evidence and provenance, and operate across customer environments without turning every deployment into a bespoke fork.
That requires more than a model endpoint and a vector database.
It requires a real platform.
Facthory is being built as that platform: a sovereign, governed, resilient operational intelligence layer for enterprises that expect AI to become part of how the organization actually runs.