Cost & Budget Model

This page follows the CSCoE documentation template. The repository identifies cost drivers and configured capacity, but does not contain verified prices, billing ownership, or an approved budget. TO_BE_CONFIRMED entries require information from the service owner or finance contact.

Licensing & pricing

Component Cost driver Rate or agreement

Model generation

Input and output tokens for the selected model; direct OpenAI or the configured Galileo/Portkey route

Contracted rate, gateway charges, and license holder: TO_BE_CONFIRMED

Application hosting

Kubernetes compute, ingress/network traffic, container registry, persistent storage, and operational services

Hosting allocation and internal rate: TO_BE_CONFIRMED

Slack and GitHub

Existing workspace/repository subscriptions and any application-related charges

Incremental licensing cost and contract owner: TO_BE_CONFIRMED

Confluence and optional Jira

Existing product licensing and any MuleSoft/MCP integration charges

Incremental licensing cost and contract owner: TO_BE_CONFIRMED

Documentation hosting

Cloudflare Pages builds/hosting and the organization’s deployment arrangement

Plan, included allowance, and any incremental charges: TO_BE_CONFIRMED

Do not assume an integration or shared platform is free because no price appears in this repository. Use the contracted model/gateway rate rather than public list pricing when estimating internal spend.

Who pays, and how

  • Funding model (chargeback, showback, or central funding): TO_BE_CONFIRMED.

  • Budget owner and billing contact: TO_BE_CONFIRMED.

  • Cost center and contract/license holder: TO_BE_CONFIRMED.

  • Process and approver for adding a consuming team: TO_BE_CONFIRMED.

Budget & usage limits

  • Approved budget, currency, and reporting period: TO_BE_CONFIRMED.

  • Model/gateway quota, rate limit, and spending cap: TO_BE_CONFIRMED.

  • Alert thresholds and action when a budget or quota is reached: TO_BE_CONFIRMED.

  • Contact for more budget or capacity: TO_BE_CONFIRMED.

The application defaults to at most five thread links per command and 2,000 messages per thread. These are input bounds, not token budgets or monetary spending limits. The repository does not implement a monetary spending cap.

The checked-in Helm defaults request one replica with 100m CPU and 128Mi memory, limit it to 250m CPU and 256Mi memory, and request a 1Gi persistent volume. These are configured capacity values, not measurements of utilization or a monthly cost estimate. Environment overrides and the live cluster’s allocations must be checked before sizing a budget.

Estimating and reviewing spend

Estimate model spend for the reporting period from measured input and output tokens at the contracted rate, then add gateway, platform, storage/network, licensing, and documentation-hosting allocations. Include retries and regenerated drafts; cancelled or unpublished drafts can still incur model charges. Human approval controls publishing, not the preceding model request.

PORTKEY_APPLICATION_NAME labels model requests with application metadata; the configured Helm value is slack-kb-agent. Confirm that the gateway’s billing reports use that metadata before relying on it for cost attribution. The finance/service owner should confirm rates, actual usage, funding model, and escalation contacts before publishing a numeric estimate.