Context
The client is a Singapore-based energy company that wants to bring AI into its everyday work — natural-language equipment inspection Q&A, automated report generation, intelligent retrieval on engineering knowledge bases. The team knows model selection and how to wire AI into the business. What they don’t want to do: procure GPUs, rack them, run the data centre, write a scheduling platform, or take the night shift.
They needed the whole programme handled: GPUs run for us, model integration into business systems delivered, 24×7 someone on watch.
What we did
TWO TWO owns the full lifecycle — from hardware through model integration to daily operations:
1. Compute deployment
8 compute nodes online — GPU model and capacity sized to the client’s business profile. Rack move-in, power, network, monitoring and spares are all handled by TWO TWO.
2. Compute scheduling
A scheduling platform customised for the client:
- Resource isolation between training and inference workloads
- Priority / queue / preemption policies tiered by business role
- Usage, cost and SLA dashboards for visibility
3. Model integration with business systems
Two production models landed inside the client’s business:
- GLM — bilingual (Chinese / English) general capability, used for engineering report drafting and internal knowledge-base Q&A
- QWEN — long-context and multi-turn dialogue, used for intelligent assistants inside business systems
Both are exposed to the client’s existing tickets, knowledge base and reporting via API. Business users call AI from tools they already use.
4. 24×7 managed operations
TWO TWO’s on-call team owns node health monitoring, model service availability, performance sweeps and incident response. The client team focuses on model iteration and business integration — not on hardware and ops.
Results
- Go-live from “quarters” to “weeks” — hardware, scheduling and model integration under one delivery. The client team is on business from day one
- Cost controlled — 8 nodes sized to actual demand, avoiding the “buy a batch first, figure out later” waste
- Model integration is not siloed — GLM / QWEN reach existing tickets, knowledge base and reports via API. Business users don’t learn a new tool
- 24×7 coverage — client team can go home. They’re no longer woken up by GPU alerts
The full case is NDA-bound. Reach out and we’ll walk you through it in private.