Know what caused the spend.
Break usage down by app, model, feature, environment, user, trace, and session.
Instrument your AI applications with exact token, cost, latency, and feature attribution—without proxying traffic or surrendering provider credentials.
One server-side SDK captures provider-reported usage and the business context that a machine-level command cannot reliably infer.
Create your account →Create a Tokly workspace and app, then save the server-only write key shown once.
Add the lightweight TypeScript package to the application that calls your model provider.
npm install tokly-sdkSend provider usage through a typed adapter and attach the feature and environment that caused it.
tokly.capture(fromOpenAIResponse({ model, usage }));Inspect exact tokens, attributed cost, latency, pricing coverage, and budget thresholds in Tokly.
Copy the complete server-side example, add your app key as an environment variable, and capture the usage returned by your provider.
Pass the model and usage from an OpenAI Responses result.
import { createTokly } from "tokly-sdk";
import { fromOpenAIResponse } from "tokly-sdk/adapters";
const tokly = createTokly({
apiKey: process.env.TOKLY_API_KEY!,
});
tokly.capture(fromOpenAIResponse({
model: response.model,
usage: response.usage,
dimensions: {
environment: "production",
feature: "assistant",
},
}));
await tokly.flush();Run Tokly only on your server. Keep TOKLY_API_KEY in an environment variable and flush before a serverless request ends.
Break usage down by app, model, feature, environment, user, trace, and session.
Effective-dated model rates preserve the cost calculation that applied when each call ran.
Threshold alerts reach your team before a monthly AI bill becomes an incident.