Selection criteria, knockout questions and a weighted scoring template for executives, CIOs, heads of IT and AI leads.
In brief: Do not choose an AI platform for business based on the most impressive demo or the model currently leading a benchmark. The key question is whether the platform can become your own extensible AI stack—with controlled identity, documented data routes, replaceable models, permission-aware company knowledge, open integrations and a robust operating model.
A risky decision in enterprise AI is not necessarily a poor model response. It can also be a good point solution that cannot connect to the next use cases. The organisation may then accumulate separate identities, duplicate knowledge indexes, conflicting policies and several commercial models. What began as a rapid pilot can become the next form of shadow IT.
This guide changes the perspective: it evaluates not only today’s tool but the ability to build on it in a controlled way over the coming years.
Tool, platform or AI stack?
The terms are often used interchangeably, but they describe three different layers:
| Layer | Purpose | Typical outcome |
|---|---|---|
| AI tool | Solve a bounded task | Chat, transcription, translation or image generation |
| AI platform | Provide shared technical services for several use cases | Identity, models, knowledge, agents, APIs and logging |
| AI stack | Link technology, organisation and continuous development | Target architecture, operating model, governance, training, use-case portfolio and roadmap |
An organisation can start with a tool. Before rolling it out, however, it should determine whether that tool can later become part of the platform—or will remain another island.
Target state: the layers of a long-term enterprise AI capability
A robust target architecture can be described in seven layers:
- User access: chat, Office, mobile use, business applications and APIs
- Identity and authorisation: single sign-on, groups, roles and delegated permissions
- Model layer: approved providers, regions, routing and fallbacks
- Knowledge layer: documents, metadata, vector indexes and source permissions
- Action layer: MCP, REST, business systems, agents and deterministic workflows
- Control layer: audit, security, quality, cost and incident processes
- Operations and enablement: owners, support, training, policy and roadmap
The central test question is therefore: Can the proposed solution grow into this stack without rebuilding identity, data and integrations for each extension?
Before scoring: five knockout criteria
A high total score must not hide a fundamental gap. Decide in advance which answers disqualify a solution. For many organisations, that list includes at least the following:
- The processing route for prompts, responses, files and logs is not fully documented.
- Single sign-on and role- or group-based approvals are unavailable.
- Data, prompt templates, agent definitions or knowledge stores cannot be exported in documented formats.
- Agents with write access can only operate through an overprivileged shared technical account.
- Responsibility for updates, security incidents, model approvals and support remains unclear.
Knockout criteria are organisation-specific. Regulated environments may add stricter requirements for regions, encryption, logging or human approval.
The ten selection criteria
1. Architectural control and deployment model
Start by establishing where the platform runs and who administers each layer. “Private instance”, “dedicated” and “in your own tenant” do not mean the same thing. The relevant components include the cloud account, network, keys, databases, backups, deployment pipeline and administrative emergency access.
Test question: Which components run in our environment, which run with the operator, who has administrative access, and what remains in place if we change operating partner?
A strong answer includes an architecture diagram, a responsibility matrix and a documented handover scenario—not merely a location name.
2. Identity and end-to-end permissions
An enterprise AI platform must connect to the existing identity and access management system. Single sign-on is only the beginning. Roles and groups should reach models, agents, knowledge spaces and tools.
Test question: Is the user’s identity passed through to knowledge queries and actions in business systems, or do all integrations operate through shared service accounts?
For write actions, also define when human approval is required and how it is recorded.
3. Data control and documented processing routes
“EU hosting” describes only one part of processing. A complete data-flow diagram includes input, output, uploaded files, embeddings, search indexes, telemetry, security logs, backups and connected external services.
Test question: Can we enforce and demonstrate which model endpoints, regions, stores and tools each data class is allowed to use?
A robust platform enforces rules technically. A policy that exists only on the intranet cannot prevent a request from being routed incorrectly.
4. Model choice, abstraction and fallbacks
Multiple models are not a goal in themselves. The strategic value comes from decoupling the model layer from the interface, knowledge and integrations. A model can then be replaced based on measured quality, availability, data route or cost without rebuilding the whole use case.
Test question: What changes are required to add a new model endpoint, block a model or switch to an approved fallback in a controlled way during an outage?
Also test whether prompts, structured output and tool calls work across several providers. API compatibility alone does not guarantee equivalent output quality.
5. Company knowledge, permissions and lifecycle
With retrieval-augmented generation (RAG), finding a document is only part of the job. Source permissions, freshness, metadata, deletion, versioning and traceability of cited passages matter just as much.
Test question: What happens in the AI index when a SharePoint permission is revoked, a file is replaced or a retention period expires?
Require a test with real permission changes. A prepared demo using public sample documents does not test this critical layer.
6. Integrations, agents and workflows
Enterprise value emerges when AI works with existing systems. Distinguish three mechanisms: read-only knowledge access, open-ended agentic actions and deterministic workflows.
Test question: Does the platform support open interfaces such as REST and MCP, and can a clearly defined process be handed to a workflow engine rather than being entrusted entirely to a language model?
A useful rule: open-ended tasks belong with an agent; stable, auditable procedures belong in a workflow. Write actions need limited permissions, confirmation and error handling.
7. Security, governance and evidence
A platform does not automatically make an organisation compliant with the GDPR or EU AI Act. It should support the organisational measures technically: approvals, purpose limitation, roles, audit information, deletion and human-in-the-loop controls where required.
Test question: Can we trace who used which approved agent and model without creating a new sensitive data store through excessive logging?
Also assess prompt-injection protections, secret management, network boundaries, vulnerability management and incident procedures. Certifications provide evidence for a defined scope; they do not replace assessment of the organisation’s use case.
8. Observability, quality and AI FinOps
User counts alone do not show whether a platform creates value. Technical and business signals are needed: tool-call success, latency, errors, model consumption, cost and a defined quality measure for each use case.
Test question: Can usage and quality be attributed to a team, use case or agent, and can limits or budgets be enforced technically where needed?
Separate reporting from control. A dashboard shows the past; routing, quotas and budgets change future behaviour.
9. Adoption, AI literacy and business ownership
A technically sound system without adoption is not a success. Organisations need accessible entry points, concrete templates, trained employees and business owners capable of judging outcomes. The current Article 4 of the EU AI Act requires providers and deployers of AI systems to take measures that support the development of AI literacy among their staff and others operating or using the systems on their behalf. Relevant knowledge, experience, education and training, the context of use and the affected groups must be considered; under the current wording, no specific level of AI literacy for an individual has to be guaranteed.
Test question: Who owns content, quality and approval for each use case, and how is the necessary AI literacy built and documented?
General prompt training is not enough for an agentic procurement process. Training must grow with the risk and the actual workflow.
10. Roadmap, modularity and exit
A long-term platform must combine two apparent opposites: stable foundations and replaceable components. Identity, data rules and audit should remain reliable; models, agents and integrations need to evolve quickly.
Test question: Which components can we continue to run ourselves or with another partner, which data and configurations do we receive back, and how is the roadmap prioritised jointly?
Exit is not only a contract clause. It reveals whether the architecture truly serves your organisation or is merely accessible during the contract term.
Weighted scorecard for your requirements document
The distribution below is a starting point, not a universal truth. Adjust the weights before the first vendor meeting; otherwise the best demonstration can unconsciously influence the criteria.
| Criterion | Example weight | Evidence |
|---|---|---|
| Architectural control and deployment | 12% | Architecture diagram, RACI, handover scenario |
| Identity and permissions | 10% | Live test with two roles and a revoked permission |
| Data control and processing routes | 14% | Data-flow diagram, contracts, region configuration |
| Model choice and fallbacks | 10% | Endpoint replacement in a test environment |
| Company knowledge and lifecycle | 12% | Synchronisation and permission test |
| Integrations, agents and workflows | 12% | End-to-end demo with your sample application |
| Security, governance and evidence | 12% | Security concept, audit and incident example |
| Observability and AI FinOps | 8% | Usage, quality and budget view |
| Adoption and ownership | 5% | Training and rollout plan |
| Roadmap, modularity and exit | 5% | Roadmap process, export and operating handover |
| Total | 100% |
Scoring scale:
- 0—not available: the requirement is not met.
- 1—partial: manual, planned or subject to a substantial limitation.
- 2—met: demonstrated within the proposed scope.
- 3—strongly met: demonstrated, automatable and transferable to your target state.
Calculate score × weight for each row. A “3” without credible evidence should receive no more than a “1”. The knockout criteria defined above remain in force regardless of the total score.
Requirements you can copy into an RFP
Use the following block as a starting point and adapt it to your risk classification:
- The solution MUST document every storage and processing route for prompts, responses, files, indexes, backups and logs.
- The solution MUST support single sign-on and role- or group-based approval for models, agents, knowledge spaces and tools.
- The solution MUST allow model endpoints to be replaced without rebuilding user accounts, the knowledge base and integrations.
- The solution MUST propagate source-permission changes and deletion into connected knowledge stores.
- Write actions MUST be restrictable according to least privilege, auditable and subject to confirmation where required.
- Data, agent definitions, templates and configuration MUST be exportable in documented formats.
- A responsibility matrix MUST cover updates, security incidents, model approvals, support and operational handover.
- The solution SHOULD support open interfaces such as REST and MCP and hand deterministic procedures to a workflow engine.
How CompanyGPT maps to this architecture
At innFactory, CompanyGPT is not an isolated chatbot. It is the core on which we build the customer’s own AI stack together. The starting point is standardised; the roadmap comes from your use cases, systems and data classes.
| Requirement | Component in the CompanyGPT stack |
|---|---|
| Shared, controlled access to AI | CompanyGPT with single sign-on, roles, models and agents |
| Company knowledge and large document collections | companyRAG and companyFILES, including the permission model |
| Work in Word, Excel, PowerPoint and Outlook | companyM365 with a native Office add-in |
| Business systems and tools | MCP integrations, REST and custom connections |
| Deterministic automation | n8n as a complementary workflow layer |
| Usage and agent activity | companyDASHBOARD |
| Cross-provider routing, budgets and fallbacks | optional, standalone AI Gateway |
| Organisational adoption | AI policy, employee training, ownership model and joint expansion planning |
Not every organisation needs every component on day one. That is the point of modular architecture: the foundation is productive while knowledge, integrations and control grow according to a prioritised backlog. The organisation does not create a new tool for each use case; it extends an existing capability.
Frequently asked questions
What distinguishes an AI platform from an AI tool?
An AI tool solves a bounded task. A platform provides shared identities, data rules, models, knowledge sources, integrations and governance for many use cases. An AI stack adds operations, responsibilities, training and a long-term roadmap to that technical foundation.
Do businesses need several AI models?
Not necessarily on day one. What matters is the ability to add or replace model endpoints later without rebuilding identity, knowledge and integrations. Many organisations start with one default model and add specialist and fallback models based on measured quality, data requirements and cost.
Is hosting in an EU region enough for GDPR?
An EU region is only one criterion. The processing route for every request, contracts, subprocessors, administrative access, deletion, logging, encryption and data sent to connected tools also matter. The assessment must be made per use case and does not constitute legal advice.
How should we score vendors?
Define knockout criteria first. Score the remaining criteria from 0 to 3 and multiply each score by your weighting. Require verifiable evidence in a demonstration, documentation or a contractual schedule for every high score. This keeps the most convincing presentation from winning automatically.
How does CompanyGPT fit into a long-term AI stack?
CompanyGPT forms the platform core for user access, models, knowledge and agents in the selected customer environment. Together with the organisation, we add companyRAG, companyFILES, companyM365, n8n, companyDASHBOARD and, where needed, the separate AI Gateway along the roadmap.
Conclusion: do not buy a roadmap you cannot influence
The right enterprise AI platform does more than answer today’s prompts. It creates reusable foundations for the next knowledge sources, agents, models and processes. Architectural control, the operating model, permissions and the roadmap therefore belong in the evaluation alongside features and user experience.
If you want to test your requirements against a long-term target state rather than a product demo, we can define the appropriate weights, knockout criteria and first expansion sequence together in an architecture conversation. Our AI strategy consulting page provides the broader strategic context.
For the next steps, continue with our ChatGPT for business rollout guide and our method for selecting Claude, GPT and other models.
