Cost & Budget Model
This page follows the CSCoE documentation template.
The repository identifies cost drivers and configured capacity, but does not contain verified prices, billing ownership, or an approved budget.
TO_BE_CONFIRMED entries require information from the service owner or finance contact.
Licensing & pricing
| Component | Cost driver | Rate or agreement |
|---|---|---|
Model generation |
Input and output tokens for the selected model; direct OpenAI or the configured Galileo/Portkey route |
Contracted rate, gateway charges, and license holder: |
Application hosting |
Kubernetes compute, ingress/network traffic, container registry, persistent storage, and operational services |
Hosting allocation and internal rate: |
Slack and GitHub |
Existing workspace/repository subscriptions and any application-related charges |
Incremental licensing cost and contract owner: |
Confluence and optional Jira |
Existing product licensing and any MuleSoft/MCP integration charges |
Incremental licensing cost and contract owner: |
Documentation hosting |
Cloudflare Pages builds/hosting and the organization’s deployment arrangement |
Plan, included allowance, and any incremental charges: |
Do not assume an integration or shared platform is free because no price appears in this repository. Use the contracted model/gateway rate rather than public list pricing when estimating internal spend.
Who pays, and how
-
Funding model (chargeback, showback, or central funding):
TO_BE_CONFIRMED. -
Budget owner and billing contact:
TO_BE_CONFIRMED. -
Cost center and contract/license holder:
TO_BE_CONFIRMED. -
Process and approver for adding a consuming team:
TO_BE_CONFIRMED.
Budget & usage limits
-
Approved budget, currency, and reporting period:
TO_BE_CONFIRMED. -
Model/gateway quota, rate limit, and spending cap:
TO_BE_CONFIRMED. -
Alert thresholds and action when a budget or quota is reached:
TO_BE_CONFIRMED. -
Contact for more budget or capacity:
TO_BE_CONFIRMED.
The application defaults to at most five thread links per command and 2,000 messages per thread. These are input bounds, not token budgets or monetary spending limits. The repository does not implement a monetary spending cap.
The checked-in Helm defaults request one replica with 100m CPU and 128Mi memory, limit it to 250m CPU and 256Mi memory, and request a 1Gi persistent volume. These are configured capacity values, not measurements of utilization or a monthly cost estimate. Environment overrides and the live cluster’s allocations must be checked before sizing a budget.
Estimating and reviewing spend
Estimate model spend for the reporting period from measured input and output tokens at the contracted rate, then add gateway, platform, storage/network, licensing, and documentation-hosting allocations. Include retries and regenerated drafts; cancelled or unpublished drafts can still incur model charges. Human approval controls publishing, not the preceding model request.
PORTKEY_APPLICATION_NAME labels model requests with application metadata; the configured Helm value is slack-kb-agent.
Confirm that the gateway’s billing reports use that metadata before relying on it for cost attribution.
The finance/service owner should confirm rates, actual usage, funding model, and escalation contacts before publishing a numeric estimate.