# Axio Intelligence > Axio Intelligence is a veteran-founded US engineering firm in San Antonio, Texas. Established companies bring us in to build what they are launching next: go-to-market platforms for new product lines and divisions, data and AI platforms, cloud infrastructure as code on AWS, Google Cloud and Azure, and connected hardware — firmware, cellular and satellite connectivity, device-to-cloud platforms and edge AI. We also cut cloud costs, modernize legacy systems, run DevOps/SRE, and provide fractional CTO leadership. Website: https://axiointelligence.com/ · Contact: contact@axiointelligence.com · Book a scoping call: https://axiointelligence.com/book/ Headquarters: San Antonio, TX, United States. Founder: Brice Ayres, U.S. Army veteran. While serving in the U.S. Army, Brice built a telecom company that gave away more than one million minutes of talk time to troops deployed in Iraq and Afghanistan. Short map of the site: https://axiointelligence.com/llms.txt ## Company FAQs **What does Axio Intelligence do?** Axio Intelligence is a US engineering firm in San Antonio, Texas. We build go-to-market platforms for new product lines, data and AI platforms, and cloud infrastructure as code on AWS, Google Cloud and Azure; we build connected hardware end to end — firmware, cellular and satellite connectivity, device-to-cloud and edge AI; and we cut cloud costs, modernize legacy systems — including moving COBOL and RPG workloads off IBM mainframes and IBM i, and Oracle and SQL Server databases to PostgreSQL or MySQL — run DevOps/SRE and provide fractional CTO leadership. **Can Axio build the platform for a new product line or division?** Yes — it is one of the most common reasons established companies hire us. We build the product site, customer and dealer portals, commerce and billing, CRM and ERP integrations, and the cloud infrastructure for a new line, so it can launch without pulling the core engineering team off the roadmap. **Which clouds do you work with?** AWS, Google Cloud (GCP) and Microsoft Azure. We build and support infrastructure as code with Terraform, OpenTofu, Pulumi, AWS CDK or Bicep, including landing zones, Kubernetes (EKS, GKE, AKS), CI/CD and ongoing IaC support. **Do you work with IBM mainframes or IBM i (AS/400)?** Yes. For US manufacturers, distributors, wholesalers and mid-size retailers still running on IBM Z (z/OS and COBOL) or IBM i (AS/400 and RPG), we move COBOL and RPG workloads to AWS, Azure or Google Cloud in a modern language, starting with the batch and extract jobs that cost the most — inventory, replenishment, pricing feeds, reports. It starts with a diagnostic of the bill and the job list, and there is no upfront fee: we are paid a percentage of the net savings in licensing and compute spend. **Who hires Axio Intelligence?** Mostly established companies launching something new — a product line, a division, an AI or data capability, or a connected product — plus IoT connectivity providers who need an implementation partner, energy and DER device makers, defense and dual-use hardware teams, and industrial operators. We also built the AI data platform behind MLtwist. **Can you build firmware and cloud software for our IoT device?** Yes. We build the embedded firmware, the connectivity path, the device-to-cloud pipeline, command-and-control APIs, secure OTA updates and the fleet console, and we can design and manufacture the hardware too. One team is accountable from the board to the dashboard. **How do engagements work and how are they priced?** Most work starts with a short scoping sprint, followed by a fixed-scope, fixed-fee build with a written plan and dates. After launch, systems can move to a monthly support retainer. Fractional CTO work is a monthly retainer for a set number of days. Mainframe and IBM i migration is the exception: no upfront fee, priced as a percentage of the net savings we deliver. Database migrations can be either a fixed fee or a share of the license savings, depending on the job. **Do you work with our existing connectivity provider?** Yes. We build on the provider you already use, which is why connectivity providers refer their customers to us as an implementation partner. **Where is Axio Intelligence located?** Axio Intelligence is headquartered in San Antonio, Texas, and works with clients across the United States. Our engineering team is US-based. # Services ## New product line platforms URL: https://axiointelligence.com/services/product-platforms/ Axio Intelligence builds the complete software platform behind a new product line or division — product site, customer and dealer portals, quoting, ordering and subscription billing, CRM and ERP integrations, analytics, and the cloud infrastructure underneath — so an established company can launch a new business without pulling its core engineering team off the roadmap or hiring a department first. ### Signs you need this - A new division has a launch date and no engineering team of its own yet. - Your core engineers can’t be pulled off the main product to build it. - The new line needs its own portal, pricing, billing and support workflows — wired into the systems you already run. - Leadership wants a real product in market this year, not after a year of hiring. ### What we deliver - Product site & launch: Positioning-ready marketing site, product pages, lead capture and analytics, built to rank and to convert. - Customer, partner & dealer portals: Accounts, roles and SSO; ordering, service requests, documentation and device or subscription management. - Commerce & billing: Quoting, checkout, subscriptions and usage-based billing, with tax, invoicing and revenue reporting. - Systems integration: CRM, ERP, support desk and data warehouse integrations so the new line runs on your existing back office. - Cloud foundation: Isolated environments as code on AWS, Google Cloud or Azure, with CI/CD, monitoring and security baselines. - Handoff or run: Documentation, runbooks and hiring support for the team that takes it over — or ongoing support from us. Stack: Next.js, Elixir / Phoenix, Kafka, React, React Native, Node.js, Go, Postgres, ClickHouse, Keycloak (SSO / SAML / OIDC), OpenTelemetry, CRM / ERP / billing integrations, AWS · GCP · Azure Engagement: Discovery sprint, then a phased build to launch Timeline: First release typically 8–16 weeks ### FAQs **Can Axio build the platform for a new product line or division?** Yes. Established companies bring us in to build everything a new line needs to go to market — product site, portals, commerce and billing, integrations with the CRM and ERP you already run, and the cloud infrastructure — as a fixed-scope, phased build. **What happens after launch?** Either we hand off to the team you hire, with documentation, runbooks and help interviewing, or we keep running the platform on a monthly support retainer. You own the code either way. **Will this pull time from our existing engineering team?** Minimal. We need access, a few decisions from product owners and integration points into your systems. The build itself runs on our team. ## Data & AI platforms URL: https://axiointelligence.com/services/data-ai-platforms/ Axio Intelligence builds data and AI platforms — ingestion and transformation pipelines, human-in-the-loop labeling workflows, workforce and project management, automated quality control, and Kubernetes-orchestrated containerized workloads — that turn raw multimodal data (video, audio, images, text and 3D) into model-ready datasets. We engineered the platform MLtwist uses to deliver AI data work, including its U.S. TSA contract. ### Signs you need this - Data preparation, not modeling, is what’s slowing your AI program down. - Labeling runs on spreadsheets, email threads and one-off scripts. - Every new customer or dataset needs a custom pipeline built from scratch. - A government or enterprise buyer wants proof of data lineage, quality and security. ### What we deliver - Pipeline orchestration: Custom pipelines triggered per project, running as containerized jobs on Kubernetes that scale with the workload. - Labeling workflows: Task setup, assignment, review loops and connectors to third-party annotation tools. - Project & workforce management: Customers, projects, labelers and reviewers managed in one system, with throughput and quality visibility. - Automated quality control: Validation rules and QC on human-in-the-loop revisions before anything ships. - Versioning & lineage: Tracked, versioned data in a secure environment, with an audit trail from raw input to delivered dataset. - Delivery in any format: Transformation and packaging to the customer’s schema — JSON, standard formats such as DICOS, or custom. Stack: ClickHouse, Kafka / Redpanda, Apache Flink, Elasticsearch / OpenSearch, Apache Iceberg, Parquet, DuckDB, PostGIS / H3, Kubernetes, Python, Postgres, Object storage (S3 / GCS), Human-in-the-loop labeling, Multimodal data, DICOS Engagement: Pipeline sprint or full platform build Timeline: Pipeline sprint: 2–4 weeks · Platform: 3–6 months ### FAQs **Can Axio build a data labeling or AI data platform?** Yes. We built the platform MLtwist runs on: project and workforce management, labeling projects, automated QC, and custom pipelines triggered on Kubernetes as containerized workloads. It is the foundation MLtwist delivered its $590K U.S. TSA multimodal labeling contract on. **Did Axio work directly for the TSA?** No. MLtwist holds the TSA contract. Axio engineered the MLtwist platform that the work is delivered on. **What data types can these pipelines handle?** Video, audio, images, text and 3D data, including security-imaging formats such as DICOS — preprocessed, labeled, quality-checked and packaged in the format the model team needs. ## Performance & legacy modernization URL: https://axiointelligence.com/services/legacy-modernization/ Axio Intelligence makes slow or aging systems fast and maintainable again — profiling and fixing performance bottlenecks, tuning databases and APIs, migrating platforms, and rewriting legacy components incrementally behind stable interfaces so the business never has to stop shipping. ### Signs you need this - Pages, APIs or reports got slow as data grew, and nobody owns performance. - A legacy system blocks new features and only one person understands it. - You need to migrate clouds or off a platform without a customer-facing outage. ### What we deliver - Performance profiling: Measure first: tracing, query analysis and load tests against realistic data. - Targeted fixes: Indexing, caching, query rewrites, async processing and right-sized infrastructure. - Incremental rewrites: Strangler-pattern replacement, one bounded component at a time, with parity tests. - Migrations: Cloud-to-cloud and platform exits run as small, reversible steps. Stack: Rust, Go, Elixir / Phoenix, Postgres, Ruby on Rails, Node.js, Python, Kafka, ClickHouse, OpenTelemetry Engagement: Performance audit, then fixed-scope phases Timeline: Audit: 1–2 weeks · Phases: 4–12 weeks each ### FAQs **Do we have to freeze features during a rewrite?** No. We replace components incrementally behind stable interfaces, with parity tests, so product work continues while the old system is retired piece by piece. **Do you modernize COBOL or RPG on IBM mainframes and IBM i?** Yes, one job at a time. We move COBOL and RPG workloads off z/OS and IBM i to AWS, Azure or Google Cloud in a modern language, starting with the batch and extract jobs that cost the most, and dual-run them until the outputs match. See mainframe & IBM i migration. **Do you migrate Oracle or SQL Server to PostgreSQL?** Yes. We convert schemas and stored procedures, keep data in sync with change data capture, and cut over with a rollback path. See database migration. ## Database migration URL: https://axiointelligence.com/services/database-migration/ Axio Intelligence migrates legacy Oracle and Microsoft SQL Server databases to PostgreSQL or MySQL — self-managed, or on Amazon RDS and Aurora, Google Cloud SQL and AlloyDB, or Azure Database. We convert schemas, PL/SQL and T-SQL stored procedures and application queries, keep old and new in sync during the move, and cut over with a rollback path, so you stop paying per-core license fees without a risky big-bang weekend. ### Signs you need this - Oracle or SQL Server licensing is one of the biggest lines in the IT budget, and a renewal or audit is coming. - Business logic lives in thousands of lines of PL/SQL or T-SQL that nobody wants to touch. - You’re moving to the cloud and don’t want to carry the old license with you. - Your team already runs Postgres or MySQL elsewhere and wants one database to support. ### What we deliver - Assessment: Schema, code and query inventory: what converts cleanly, what needs rewriting, and license and hosting cost before and after. - Schema & code conversion: Tables, types, indexes, stored procedures, triggers and packages converted to PostgreSQL or MySQL, with tests for every converted routine. - Data migration & sync: Bulk load plus change data capture, so old and new stay in sync until cutover. - Cutover & rollback: Application changes, tuning against production-sized data, and a tested rollback plan. Stack: PostgreSQL, MySQL, Oracle (PL/SQL), SQL Server (T-SQL), Amazon RDS / Aurora, Cloud SQL / AlloyDB, Azure Database for PostgreSQL / MySQL, AWS DMS & Schema Conversion, Debezium (CDC), ora2pg, pgloader, Terraform / OpenTofu Engagement: Assessment, then a phased migration per database Timeline: Sized after the assessment Pricing: Fixed fee or a percentage of the license savings, depending on the job ### FAQs **Can you migrate Oracle to PostgreSQL?** Yes. We convert the schema and the PL/SQL packages, procedures and triggers to PostgreSQL, move the data with a bulk load and change data capture, and cut over once the application has run against Postgres with production-sized data. **Can you migrate SQL Server to PostgreSQL or MySQL?** Yes. T-SQL procedures, functions and triggers are converted and tested routine by routine, application queries are updated, and data stays in sync until cutover. **Do we need downtime to switch databases?** Usually only a short cutover window. Change data capture keeps the new database current while the application is tested against it, and the old database stays available as a rollback until you are confident. **PostgreSQL or MySQL — which should we choose?** PostgreSQL is usually the closer fit for Oracle and SQL Server workloads with heavy procedural code, rich types and complex queries. MySQL makes sense when your team already runs it or the workload is simpler. The assessment recommends one based on your code. **How is database migration priced?** Either a fixed fee or a percentage of the license savings, depending on the job. When the goal is getting off Oracle or SQL Server licensing, a share of the savings usually fits; when the migration is part of a wider build, a fixed fee usually does. We recommend one after the assessment. **Can you migrate Db2 too?** Yes. Db2 for z/OS and Db2 for i data moves to PostgreSQL as part of our mainframe and IBM i work, alongside the jobs that read it. ## Fractional CTO & engineering URL: https://axiointelligence.com/services/fractional-cto/ Axio Intelligence provides fractional CTO and fractional engineering leadership for hardware, IoT and defense-technology companies — owning architecture decisions, the technical roadmap, hiring, vendor and connectivity selection, and investor or customer technical diligence — part-time, until a full-time leader is justified. ### Signs you need this - You’re a hardware or domain founder without a senior software leader. - An SBIR or pilot is turning into a program and the engineering has to grow up fast. - You need someone to evaluate vendors, contractors or candidates you can’t judge yourself. ### What we deliver - Architecture ownership: Decisions written down, reviewed and defensible to customers and investors. - Roadmap & delivery: A plan the team can execute, with cadence, metrics and risk tracking. - Hiring & vendors: Role design, interviews, contractor and connectivity vendor selection. - Diligence support: Technical answers for investors, primes, utilities and security reviews. Stack: Architecture decision records, Roadmapping, Hiring loops, Vendor evaluation, Security questionnaires Engagement: Monthly retainer, a set number of days per month Timeline: Typically 3–12 months ### FAQs **What does a fractional CTO do for a hardware startup?** Owns the technical decisions a full-time CTO would — architecture, connectivity and cloud choices, security, hiring and vendor selection — for a set number of days per month, and can bring an engineering team to execute the plan. **How is this different from hiring an agency?** An agency delivers a project. A fractional CTO is accountable for your technical direction across projects and vendors, including the ones we don’t build. ## IoT product engineering URL: https://axiointelligence.com/services/iot-product-development/ Axio Intelligence builds the software layer that turns a connected device into a product you can sell: embedded firmware, device-to-cloud ingestion, command-and-control APIs, secure over-the-air updates, and the fleet console your operations team and customers use every day. We take a device from working prototype to a managed fleet in the field. ### Signs you need this - The hardware works on the bench, but there’s no reliable way to reach, update or command units in the field. - Your connectivity vendor shipped SIMs and modules, and the project stalled on firmware and cloud. - Customers or a utility want an API, a dashboard and uptime numbers before they buy. - You need OTA updates you can trust before the next thousand units ship. ### What we deliver - Embedded firmware: Device application, power budgeting, watchdogs, store-and-forward, and command handling on Zephyr, FreeRTOS or embedded Linux. - Device-to-cloud pipeline: Ingestion over MQTT, LwM2M, CoAP or HTTPS into a time-series store, with device identity and TLS end to end. - Command & control API: Acknowledged commands, desired-vs-reported state, scheduling and audit trail, so every command is traceable. - Secure OTA: Signed, staged, resumable firmware updates with rollback and fleet cohorts. - Fleet console: Operator and customer dashboards: health, alerts, configuration, maps and exports. - Integrations: ERP, CRM, utility and C2 systems via REST, webhooks and message queues. Stack: Rust, C / C++, Zephyr RTOS, FreeRTOS, Embedded Linux / Yocto, Buildroot, nRF91, Quectel, MQTT / Sparkplug B, LwM2M, CoAP, OPC UA, NATS JetStream, AWS IoT Core, Blues Notecard, Particle, Memfault, Mender / SWUpdate / RAUC, TUF / Uptane, TimescaleDB, ClickHouse Engagement: Scoping sprint, then a fixed-scope build to a fielded pilot Timeline: Typical pilot-to-production build: 8–16 weeks ### FAQs **Can you pick up an IoT project that has stalled?** Yes — that is the most common way we start. We audit what exists (hardware, firmware, cloud, connectivity contract), keep what works, and ship the missing layer, usually firmware command handling, the cloud pipeline and a fleet console. You keep the code and the IP. **Do you work with our existing connectivity provider?** Yes. We build on the connectivity you already have — cellular, satellite or private LTE — and hand connectivity decisions back to you and your provider. **Which cloud do you build on?** Usually AWS, and also GCP, Azure or your own infrastructure. We default to boring, well-supported managed services so your team can run the system after hand-off. **How do you handle over-the-air updates safely?** Signed images, staged rollouts by cohort, resumable downloads over constrained links, health checks after update, and automatic rollback. We design OTA before the first fleet ships, not after. ## Cellular & satellite connectivity URL: https://axiointelligence.com/services/cellular-satellite-connectivity/ Axio Intelligence designs and integrates the connectivity path for devices in the field — LTE-M, NB-IoT and Cat-1bis cellular, eSIM, satellite (NTN and L-band), and private LTE — including failover between links, power and data budgets, and carrier certification support. We build on the provider you already use. ### Signs you need this - Devices depend on customer Wi-Fi and drop off when a router changes. - Assets operate beyond cellular coverage and need a satellite fallback. - Data costs or battery life are higher than the business model allows. - You need a PACE-style primary/alternate path for DDIL environments. ### What we deliver - Link selection: Radio, module, SIM/eSIM and plan trade study against coverage, power, data volume and unit cost. - Cellular integration: Modem bring-up, PSM/eDRX power tuning, network registration logic and reconnection strategy. - Satellite paths: NTN NB-IoT, Iridium SBD/Certus or Globalstar messaging for off-grid and backup links. - Failover logic: Link health scoring, automatic fallback, message prioritisation and store-and-forward. - eSIM & provisioning: SGP.32 remote SIM provisioning, profile management and factory provisioning flows. - Certification support: Preparing devices and documentation for carrier and module certification. Stack: LTE-M, NB-IoT, Cat-1bis, eSIM SGP.32, 3GPP NTN, Skylo, Iridium SBD / Certus, Globalstar, Private LTE / CBRS, 900 MHz LTE, PSM / eDRX Engagement: Connectivity assessment, then integration inside a product build Timeline: Assessment: 1–2 weeks · Integration: 4–10 weeks ### FAQs **LTE-M or NB-IoT — which should our device use?** LTE-M suits devices that move, need lower latency or send more data, and it supports voice-grade handover. NB-IoT suits stationary, deep-indoor, very low-data devices. Many products ship modules that support both and choose per region; we test against real coverage before committing. **Can we add satellite as a backup to cellular?** Yes. We design a failover path — typically 3GPP NTN or Iridium — with message prioritisation, so critical commands and alarms still get through when cellular is unavailable, without paying satellite rates for routine telemetry. **Do we have to switch connectivity providers?** No. We work with the provider you choose and build on the network you already use, which is why connectivity providers refer their customers to us. ## Edge AI & computer vision URL: https://axiointelligence.com/services/edge-ai-computer-vision/ Axio Intelligence deploys machine learning where the data is created: computer vision detection and tracking on edge hardware such as NVIDIA Jetson, anomaly detection on sensor and fleet telemetry, and the data pipelines, labeling workflows and model monitoring needed to keep models accurate in the field. ### Signs you need this - Streaming raw video or telemetry to the cloud is too slow, too expensive or not possible. - You have months of sensor data and no early warning when equipment starts to fail. - A model works in a notebook but not on the device, at the power budget you have. - Field data isn’t making it back into training in a usable form. ### What we deliver - On-device inference: Model optimisation and deployment on Jetson, ARM and x86 edge compute with TensorRT or ONNX Runtime. - Computer vision: Detection, classification and tracking pipelines for cameras and EO/IR sensors. - Anomaly detection: Baselines and alerting on telemetry to catch drift, faults and tampering early. - Data & labeling pipelines: Getting field data into models: capture, quality checks, labeling workflows and dataset versioning. - Model operations: Versioned model delivery over OTA, shadow deployment and on-device performance monitoring. - Field validation: Accuracy, latency and power measured on the target hardware in real conditions, with a written report before rollout. Stack: NVIDIA Jetson, DeepStream, TensorRT, Holoscan, OpenVINO, ONNX Runtime, GStreamer, MISB KLV / STANAG 4609, ROS 2, DDS (RTI Connext / Cyclone), OpenCV, PyTorch, C++, Rust, Python Engagement: Feasibility spike on your data, then a production build Timeline: Feasibility: 2–4 weeks · Production: 6–12 weeks ### FAQs **Can you run computer vision on a drone or a pole-mounted camera?** Yes. We size the model to the compute and power available — typically NVIDIA Jetson-class hardware — and send detections and events upstream instead of raw video. **What does anomaly detection need to get started?** A few weeks of representative telemetry and a list of the failures you care about. We start with simple statistical baselines that are explainable to operators, then add learned models where they earn their keep. ## Hardware design & manufacturing URL: https://axiointelligence.com/services/hardware-manufacturing/ Axio Intelligence designs and manufactures connected hardware alongside the software that runs on it — electronics, enclosures, design for manufacturing, factory test and provisioning, and production builds — so one team is accountable from the circuit board to the cloud console. ### Signs you need this - You have a proven prototype and need a build that survives weather, vibration and installers. - Your hardware and software vendors blame each other when units fail in the field. - Factory provisioning — certificates, SIM profiles, serial numbers — is manual and error-prone. ### What we deliver - Electronics & enclosure: Schematic, PCB layout, antenna placement and ruggedised enclosures for the operating environment. - Design for manufacturing: BOM optimisation, second-source components and documented country of origin when your program requires it. - Factory test & provisioning: Test fixtures, burn-in, and automated provisioning of device identity, certificates and SIM profiles. - Production builds: Pilot runs through production quantities with traceability from serial number to firmware version. Stack: PCB design, RF & antenna integration, IP-rated enclosures, DFM / DFT, Factory provisioning, Device PKI Engagement: Design review, then prototype and pilot-run builds Timeline: Varies by design maturity — scoped after review ### FAQs **Can you build hardware and write the firmware and cloud software?** Yes. That is the point: one accountable team from the board to the console, so provisioning, OTA and field diagnostics are designed into the hardware from the start. **Can you work from our existing design?** Yes. Most engagements start from an existing prototype or production design; we review it for manufacturability, connectivity and field reliability before building. ## Cloud infrastructure & IaC URL: https://axiointelligence.com/services/cloud-infrastructure-iac/ Axio Intelligence designs, builds and supports cloud infrastructure on AWS, Google Cloud (GCP) and Azure, written as code in Terraform, OpenTofu, Pulumi or native tooling — landing zones and account structure, networking, identity and security baselines, Kubernetes platforms (EKS, GKE, AKS), CI/CD and GitOps — and provides ongoing Infrastructure as Code support for teams whose infrastructure has drifted from its code. ### Signs you need this - Infrastructure was clicked together in the console and nobody can reproduce it. - Terraform exists, but plans are scary and the state has drifted. - A new product line needs its own isolated environments — quickly and safely. - You run on more than one cloud, or a customer requires one you don’t use yet. ### What we deliver - Landing zones & accounts: AWS Organizations / Control Tower, GCP organizations and folders, Azure management groups — with guardrails. - Terraform / OpenTofu: Reusable modules, remote state, environments and a review workflow your team trusts. - Network, identity & security: VPC/VNet design, private connectivity, IAM and SSO, secrets, logging and baseline policies. - Kubernetes platforms: EKS, GKE or AKS clusters with ingress, autoscaling, observability and upgrade paths. - CI/CD & GitOps: Plan/apply pipelines, policy checks, drift detection and Argo CD or Flux for workloads. - IaC support: Monthly support for reviews, provider upgrades, drift cleanup and new environments. Stack: AWS · GCP · Azure, Terraform, Kubernetes, OpenTofu, Pulumi, AWS CDK, CloudFormation, Bicep, EKS · GKE · AKS, K3s · RKE2, Zarf (air-gapped), Helm, Argo CD, HashiCorp Vault, GitHub Actions, AWS GovCloud Engagement: Foundation build, then optional monthly IaC support Timeline: Foundation: 3–6 weeks · Support: month to month ### FAQs **Which clouds does Axio support?** AWS, Google Cloud (GCP) and Microsoft Azure, including multi-cloud setups. We write the infrastructure as code — usually Terraform or OpenTofu, or Pulumi, AWS CDK or Bicep where a team already uses them. **Can you take over an existing Terraform codebase?** Yes. We start by reconciling state with what is actually deployed, then refactor into modules, add a plan/apply pipeline and document it — without a big-bang rewrite. **Do you offer ongoing Infrastructure as Code support?** Yes. A monthly IaC support retainer covers change reviews, provider and Kubernetes upgrades, drift cleanup, new environments and on-call escalation. ## Cloud cost reduction URL: https://axiointelligence.com/services/cloud-cost-reduction/ Axio Intelligence reduces cloud spend on AWS, GCP and Azure. A read-only diagnostic inventories every resource, maps spend to workloads and quantifies what is recoverable; a senior engineer validates the findings and presents a prioritised, fixed-fee plan. Published optimisation benchmarks commonly put recoverable spend at 20–50% for teams that have never done this work. ### Signs you need this - The bill grew faster than revenue and nobody can say exactly why. - IoT telemetry and video storage costs are climbing with every device you ship. - Finance is asking engineering questions engineering can’t answer quickly. ### What we deliver - Read-only diagnostic: Audit-scoped access, full resource inventory, dependency map and spend attribution. - Findings deck: Quantified opportunities ranked by dollars saved per week of work — yours to keep. - Fixed-fee execution: Rightsizing, commitment strategy, storage tiering, data-transfer fixes and architecture changes. - Guardrails: Budgets, tagging policy and alerts so spend doesn’t drift back. Stack: AWS, GCP, Azure, Kubernetes, Terraform / OpenTofu, Savings Plans, S3 / object storage tiering Engagement: Diagnostic, then optional fixed-fee execution Timeline: Diagnostic: ~10 days · Execution: 4–8 weeks ### FAQs **Is the diagnostic safe to run against production?** Yes. It uses read-only, audit-scoped credentials you can revoke at any time. Nothing is changed until you approve a plan. **How much can we expect to save?** It depends on the workload. Published benchmarks from the cloud providers and the FinOps Foundation commonly show 20–50% recoverable for teams that haven’t optimised before; the diagnostic gives you your actual number. ## Get the jobs that trap you off the IBM box. URL: https://axiointelligence.com/services/mainframe-migration/ Axio Intelligence helps US manufacturers, distributors, wholesalers and mid-size retailers cut IBM, Broadcom and BMC mainframe licensing and compute cost by moving COBOL workloads to AWS, Azure or Google Cloud in a modern language — starting with the peak batch and extract jobs that cost the most: inventory, replenishment, pricing feeds and nightly reports. Go as far as a full migration, or stop when the savings stop. There is no upfront fee: we are paid a share of the net savings. ### Signs you need this - Mainframe licensing and compute cost keeps climbing, and nobody can say which workloads drive it. - Inventory, replenishment and pricing jobs run at the busiest time of the month, and nobody wants to touch them. - On IBM i, the one person who understands the RPG that runs the business is close to retiring. - Leadership wants off the box, but not through a multi-year, all-at-once rewrite. ### What we deliver - Ranked job inventory: Every scheduled batch and extract job, with its inputs, outputs and owner, and what it costs you or risks for the business. - Migration plan & savings baseline: Which jobs move, in what order, and today’s licensing and compute spend — the baseline the savings are measured against. - Workloads rebuilt in the cloud: On AWS, Azure or Google Cloud in a modern language, with Postgres or your existing database — scheduled, monitored and documented. - Parity evidence: Dual-run reports showing old and new outputs match before anything on the IBM side is switched off. ### How it works 1. Read the bill and the job list: Contracts, invoices and usage reports, plus the job scheduler and programs. You send exports — we don’t need write access to anything. 2. Pick one job: One file-in, file-out job or extract that is expensive or risky. Start small, prove it, then widen. 3. Dual-run until outputs match: The new job runs in the cloud next to the old one, every cycle, and the outputs are compared until they match. 4. Turn the old job off: Only after the outputs match and the people downstream sign off. Then the next job, wave by wave. Stack: COBOL, z/OS · JCL, IBM i · RPG IV / ILE, CL, Db2 for z/OS · Db2 for i, VSAM, Control-M · CA-7 · IBM Z Workload Scheduler, SCRT · SMF, Postgres, Change data capture, AWS · Azure · GCP, Terraform / OpenTofu Engagement: Diagnostic, then migration in waves, one job at a time Timeline: Sized to your job list after the diagnostic Pricing: No upfront fee — a percentage of the net savings Not a fit: Bank deposit and money-movement cores, card authorization, insurance claims adjudication and government benefit systems. ### FAQs **How is mainframe migration priced?** There is no upfront fee. We are paid a percentage of the net savings: the drop in your annual IBM and third-party licensing and mainframe compute spend once workloads move, after the cost of running them in the cloud, measured against a baseline we agree with you before work starts. **Our contract isn’t based on the rolling 4-hour peak. Does this still apply?** Yes. Mainframe software is licensed many ways — rolling 4-hour peak, full capacity, consumption-based or enterprise agreements — and the details vary by vendor and contract. Whatever the model, the bill follows the work on the mainframe. Moving that work to the cloud is what lets you license and run less of it. **What is a rolling 4-hour peak?** It is a common way IBM and other vendors price z/OS software: the bill is based on the highest four-hour average of CPU use in the month, as reported by IBM’s Sub-Capacity Reporting Tool (SCRT). One busy window can set the bill for the whole month. It is one pricing model among several. **Is IBM i the same as a mainframe?** No. IBM i (formerly AS/400 and iSeries) is the operating system on IBM Power servers, licensed differently from IBM Z mainframes running z/OS. What the two share is decades of RPG or COBOL code that few people can still maintain, so we run them as separate diagnostics with separate goals. **Do we have to move everything at once?** No. We start with the batch and extract jobs that cost the most or carry the most risk, and move in waves. You can go all the way to a full migration, or stop when the remaining savings aren’t worth it. **Will you need a copy of all our data?** Not up front. Each job only needs the files or tables it reads. We replicate those inputs to the cloud — as a scheduled extract or a change data capture feed — and widen the scope as more workloads move. ## Infrastructure evaluation URL: https://axiointelligence.com/services/infrastructure-assessment/ Axio Intelligence evaluates software and cloud infrastructure independently — architecture, security posture, reliability, scalability, cost and team practices — and delivers a ranked findings report with a practical roadmap. Leaders use it before a scale-up, a fundraise, an acquisition or a major customer’s security review. ### Signs you need this - An investor, acquirer or large customer is about to look under the hood. - You’re about to 10× device count or users and aren’t sure what breaks first. - A new engineering leader needs an honest baseline in weeks, not quarters. ### What we deliver - Architecture review: Current-state diagrams, data flows and single points of failure. - Security posture: Identity, secrets, network exposure, device security and supply-chain hygiene (SBOM). - Reliability & scale: Load characteristics, bottlenecks and failure modes at the next order of magnitude. - Ranked roadmap: Findings ordered by risk and effort, with fixed-fee options for each. Stack: AWS / GCP / Azure, Kubernetes, Terraform, NIST SP 800-171 alignment, DISA STIG-aligned baselines, SBOM (SPDX / CycloneDX), SLSA, SPIFFE / SPIRE, Threat modeling Engagement: Fixed-fee assessment Timeline: Typically 2–3 weeks ### FAQs **What do we get at the end of an infrastructure evaluation?** A written report and a live walkthrough: current-state architecture, ranked findings with risk and effort, and a roadmap. It is useful on its own — you are not obligated to hire us for the fixes. ## DevOps, SRE & ongoing support URL: https://axiointelligence.com/services/devops-sre/ Axio Intelligence runs DevOps and site reliability engineering for teams that ship connected products — CI/CD for cloud services and firmware, infrastructure as code, observability, SLOs and alerting, incident response, and ongoing support on a monthly retainer after launch. ### Signs you need this - Deploys are manual, risky, or depend on one person being online. - You find out about outages from customers. - The system launched and now nobody has time to run it. ### What we deliver - CI/CD: Automated build, test and deploy for services and firmware, with signed artifacts and SBOMs. - Infrastructure as code: Terraform/OpenTofu environments on AWS, Google Cloud or Azure that are reproducible and reviewable. - Observability & SLOs: Metrics, logs, traces and device-fleet health with alerts tied to user-facing objectives. - Ongoing support: On-call, patching, capacity planning and a monthly reliability review. Stack: Kubernetes, Terraform / OpenTofu, OpenTelemetry, Prometheus, Grafana, Loki, Argo CD, GitHub Actions, Sigstore / cosign, Iron Bank hardened images, FIPS 140-3 validated crypto, HashiCorp Vault, Sentry, AWS GovCloud-ready patterns Engagement: Setup project, then a monthly support retainer Timeline: Setup: 3–8 weeks · Support: month to month ### FAQs **Do you offer ongoing support after a project ships?** Yes. Most builds move to a monthly support retainer covering monitoring, on-call, patching, OTA releases and a monthly reliability review. It is optional — we also document and hand off to your team. # Industries ## Connectivity providers & IoT platforms URL: https://axiointelligence.com/industries/iot-connectivity-providers/ Axio Intelligence is an IoT systems integrator and implementation partner for connectivity providers — cellular MVNOs, eSIM platforms, module makers, satellite IoT networks and private LTE operators. When a customer buys SIMs or modules and then stalls on firmware, cloud, command-and-control or dashboards, we build that layer on your stack, and the customer stays yours. ### Problems we solve - Accounts that stall after the SIM order: The OEM signs, orders SIMs, then waits months on firmware and cloud. Activations — and your revenue — wait with them. - Custom work your team shouldn’t staff: Solutions engineers get pulled into application builds that don’t scale across your customer base. - Churn risk when projects fail: A failed pilot looks like a connectivity problem even when it was a software problem. ### What we build - Device application & firmware: On your modules and SIMs, tuned for your network’s power and data profile. - Device-to-cloud & command path: Ingestion, acknowledged commands and OTA, integrated with your platform APIs. - Fleet console: The dashboard the OEM’s operations team and end customers actually use. - Vertical integrations: Utility, industrial, logistics and C2 integrations your customers ask for. ### Protocols and standards - Cellular & eSIM: LTE-M, NB-IoT, Cat-1bis, SGP.32 eSIM, Multi-IMSI - Satellite & private networks: 3GPP NTN, Iridium SBD / Certus, Globalstar, CBRS private LTE, 900 MHz LTE - Device-to-cloud: MQTT, LwM2M, CoAP, HTTPS / webhooks, AWS IoT Core - Platforms we build on: Blues Notecard / Notehub, Particle, Digi Remote Manager, Telit deviceWISE, MultiTech Conduit Rules of engagement: We build on your network and don’t move referred customers to another one. Connectivity decisions stay with you. We report status back to your partner team on every referred account. ### FAQs **How does a connectivity provider refer an account to Axio?** Send an introduction. We run a joint scoping call within two business days, quote a fixed-scope project to the customer, and keep your partner manager updated through delivery. The customer contracts with us for engineering and with you for connectivity. **Will Axio recommend a different carrier to our customer?** No. We build on the connectivity the customer already chose. If there is a real technical constraint, we raise it with you first. **Can Axio act as overflow for our solutions engineering team?** Yes. Many providers use us as a white-label or co-branded engineering bench for application work that falls outside their core platform. ## Energy & DER devices URL: https://axiointelligence.com/industries/energy-der/ Axio Intelligence builds the software that lets utilities, aggregators and virtual power plants dispatch your devices — a reliable cellular command path, IEEE 2030.5, OpenADR 2.0b/3.0, SunSpec Modbus and OCPP interfaces, aggregator and DERMS integrations, fleet telemetry for measurement and verification, and secure OTA — for makers of home and commercial battery storage, solar inverters and backup power systems, EV chargers and load-control equipment. ### Problems we solve - Wi-Fi-only devices can’t be relied on: Utilities now buy firm capacity and fast response. Home Wi-Fi drops, routers change, and dispatch success suffers. - A new protocol for every program: Each utility, aggregator or state rule means another adapter and another certification cycle. - Program approval blocks sales: If your device isn’t on the program’s approved list, the utility can’t buy it — no matter how good the hardware is. - Cybersecurity is becoming a gate: Signed firmware, device identity and SBOMs are showing up in utility and standards requirements. ### What we build - Cellular command path: LTE-M/Cat-1bis connectivity with acknowledged commands and store-and-forward, independent of customer Wi-Fi. - Utility & aggregator adapters: IEEE 2030.5 client, OpenADR VEN, SunSpec Modbus and aggregator API integrations. - Fleet console & M&V data: Availability, dispatch acknowledgements, load shed and baselines your program partners can audit. - Secure OTA & device identity: Signed, staged updates and per-device certificates across the installed base. ### Protocols and standards - Device & grid interfaces: IEEE 2030.5 / CSIP, OpenADR 2.0b, OpenADR 3.0, SunSpec Modbus, OCPP 1.6 / 2.0.1, BMS over CAN / Modbus - Programs & integrations: Utility DR programs, VPP aggregators, DERMS, ERCOT ADER, CA SGIP - Security & standards: UL 2941, IEEE 1547.3, NISTIR 7628, Signed OTA, SBOM - Device-to-cloud: LTE-M / Cat-1, Wi-Fi & BLE commissioning, MQTT over TLS, AWS IoT Core, Cloud-to-cloud APIs ### FAQs **How do I make my energy device dispatchable by a utility?** You need a reliable command path to every unit (usually cellular, not customer Wi-Fi), a supported interface — IEEE 2030.5, OpenADR or an aggregator API — acknowledgements and telemetry that prove the device responded, and a fleet console for your program partners. Axio builds all four. **Should we implement OpenADR 2.0b or 3.0?** Most existing utility programs still specify OpenADR 2.0b; OpenADR 3.0 is simpler to build against and gaining adoption. We usually recommend a VEN architecture that can speak both, so one integration covers current and upcoming programs. **Can you add cellular to a home battery or solar inverter?** Yes — as a gateway alongside the inverter or battery management system, or integrated into your controller — with a cloud command path your utility and VPP partners can dispatch against. ## Defense & dual-use URL: https://axiointelligence.com/industries/defense-dual-use/ Axio Intelligence builds the software around defense and dual-use hardware — drones, counter-UAS sensors, radios and SATCOM terminals: integrations with TAK/Cursor on Target and other C2 systems, ground control and operator consoles, fleet health, remote configuration and secure OTA, cellular/SATCOM failover for DDIL conditions, edge AI, and the data pipelines from field to model. Engagements start unclassified. ### Problems we solve - “Integrate once, show up everywhere”: Every sensor and effector has to appear in the operator’s common operating picture, through more than one C2 system. - Fleets outgrow the startup dashboard: After first fielding, health monitoring, configuration and OTA across units becomes the bottleneck. - Degraded comms are the default: Links drop. Systems need primary/alternate paths and graceful store-and-forward. - Prototype to program: A contract award turns a demo into a delivery schedule, faster than the software team can grow. ### What we build - C2 & COP integration: TAK plugins and CoT bridges, and connectors to C2 platforms through their published SDKs and APIs. - Operator & fleet software: Ground control, mission and operator consoles; fleet health, configuration and after-action data. - Secure OTA & device lifecycle: Signed updates for embedded Linux and Jetson-class devices, provisioning and depot tooling. - Resilient comms: Cellular, SATCOM and mesh path selection with failover and message prioritisation. - Edge AI & data pipelines: Detection and anomaly models on the edge, and the pipeline that gets field data into models. ### Protocols and standards - C2 & interoperability: TAK / ATAK plugins, Cursor on Target (CoT), DDS (RTI Connext / Cyclone), gRPC / Protobuf, Cesium / MapLibre mapping, MOSA-aligned interfaces - Vehicles & autonomy: MAVLink, PX4 / ArduPilot, ROS 2, NVIDIA Jetson / DeepStream, MISB KLV / STANAG 4609 video - Comms: LTE / private LTE, SATCOM (LEO / GEO), Iridium, PACE / DDIL design, NATS store-and-forward - Delivery: Rust / C++ on Yocto Linux, Air-gapped K8s: RKE2 + Zarf, Iron Bank hardened images, FIPS 140-3 validated crypto, Signed builds, SBOM & SLSA, NIST SP 800-171-aligned practices How we engage: Unclassified, fixed-scope integration sprints; subcontracting on SBIR/STTR and prototype awards; and teaming with primes and small businesses. We do not hold a facility clearance, and we say so up front. ### FAQs **Can Axio build a TAK plugin or Cursor on Target integration for our system?** Yes. We build TAK plugins and CoT bridges that publish your device’s tracks, status and alerts into the operator’s TAK environment, alongside REST/gRPC connectors for other C2 platforms. **Does Axio hold a security clearance?** No. Axio Intelligence does not hold a facility clearance. Our engagements are unclassified — fleet software, commercial connectivity, test telemetry, integrations and data pipelines — and we are transparent about that in scoping. **Can Axio subcontract on an SBIR or prototype award?** Yes. We work as a software and hardware subcontractor to small businesses on SBIR/STTR and prototype awards, and as a teaming partner for integration work primes don’t want to staff. ## Industrial & remote assets URL: https://axiointelligence.com/industries/industrial-remote-assets/ Axio Intelligence builds monitoring and control for assets far from the network — wellheads, pump stations, tanks, cold storage, water systems and municipal infrastructure: telemetry over cellular and satellite, Modbus and SCADA integration, anomaly detection that warns before failure, and dashboards and alerts for operations teams. We also modernize the monitoring platforms that already run these fleets. ### Problems we solve - Truck rolls to check on equipment: Someone drives out to read a gauge that could have reported itself. - Legacy monitoring platforms: A system built a decade ago is slow, expensive to run and hard to extend. - Alarms without context: Operators get noise, not early warning. ### What we build - Telemetry & control: Sensor and PLC integration over Modbus, with cellular or satellite backhaul. - Anomaly detection: Baselines per asset that flag drift and failure before it becomes downtime. - Operations dashboards: Maps, trends, alerts and reports your field teams trust. - Platform modernization: Performance, cost and reliability work on existing monitoring platforms. ### Protocols and standards - Field: Modbus RTU / TCP, OPC UA, SCADA integration, MQTT Sparkplug B, 4–20 mA / digital I/O - Links: LTE-M / Cat-1, Satellite IoT, LoRaWAN, Private LTE - Platform: ClickHouse / TimescaleDB, Kafka / Flink streaming, Anomaly detection, Alerting & escalation, PostGIS / H3 geospatial - Operations: Remote config & OTA, Store-and-forward buffering, Solar & battery power budgets, Work-order & ERP integration ### FAQs **Can you connect existing field equipment without replacing it?** Usually, yes. We add a gateway that reads existing controllers and sensors over Modbus or digital I/O and backhauls over cellular or satellite, so the installed equipment stays in place. ## Manufacturing & distribution URL: https://axiointelligence.com/industries/manufacturing-distribution/ Axio Intelligence works with US manufacturers, distributors, wholesalers and mid-size retailers that still run the business on IBM mainframes (z/OS and COBOL), IBM i (AS/400 and RPG), or legacy Oracle and SQL Server databases. We move those workloads — starting with inventory, replenishment, pricing feeds, nightly reports and partner file exchanges — to AWS, Azure or Google Cloud and PostgreSQL or MySQL, cutting licensing and compute cost and the dependence on the last person who understands the code. ### Problems we solve - Licensing that grows every year: IBM, Broadcom, BMC, Oracle or Microsoft — whatever the contract model, the bill rises with the work on the old platform. - Batch nobody wants to touch: Inventory, replenishment and pricing jobs feed the warehouse, the stores and EDI partners. If one is late, someone screams. - The last person who knows the code: The RPG, COBOL and stored procedures that run the business depend on a shrinking group of people. ### What we build - Mainframe batch migration: COBOL batch and extract jobs on z/OS moved to the cloud in a modern language, starting with the ones that cost the most. - IBM i batch & extract offload: RPG and COBOL batch and file extracts rebuilt in the cloud, documented and owned by more than one person. - Database migration: Db2, Oracle and SQL Server moved to PostgreSQL or MySQL, with data kept in sync until cutover. - Inventory, pricing & replenishment feeds: Extracts landed in Postgres or your warehouse, so planning, pricing and e-commerce read from the cloud instead of the box. ### Protocols and standards - IBM Z: z/OS, COBOL, JCL, Db2 for z/OS, VSAM - IBM i: IBM i (AS/400, iSeries), RPG IV / ILE RPG, CL, Db2 for i, Physical & logical files - Databases: Oracle → PostgreSQL, SQL Server → PostgreSQL / MySQL, Db2 → PostgreSQL, Change data capture, EDI (X12) & SFTP feeds - Where it lands: AWS · Azure · Google Cloud, PostgreSQL · MySQL, Amazon RDS / Aurora, Snowflake / BigQuery, Terraform / OpenTofu How it’s priced: Mainframe and IBM i migration has no upfront fee: we are paid a percentage of the net savings in licensing and compute spend, measured against a baseline we agree before work starts. ### FAQs **Can you move inventory and pricing feeds off the mainframe?** Yes. They are typical first candidates: they read IBM data and write a file, feed or report, so they can be rebuilt in the cloud and run side by side with the old job until the outputs match. **Can you migrate our Oracle or SQL Server databases too?** Yes. We migrate Oracle, SQL Server and Db2 to PostgreSQL or MySQL, converting schemas and stored procedures and keeping data in sync until cutover. **Who is this not for?** Bank deposit and money-movement cores, card authorization, insurance claims adjudication and government benefit systems. Our focus is manufacturers, distributors, wholesalers and mid-size retailers. # Case studies ## The AI data platform behind MLtwist URL: https://axiointelligence.com/work/mltwist-ai-data-platform/ Client: MLtwist · Sector: AI data platforms · Status: In production We designed and built the platform MLtwist runs its AI data business on — projects, people, labeling and Kubernetes pipelines. Our role: Axio designed and built the MLtwist platform. Challenge: MLtwist turns raw, messy data into model-ready datasets for government, research and enterprise customers. Every customer brings different data types, labeling tools and delivery formats, so the business needed one platform to run projects, people and pipelines — without custom engineering for every engagement. ### Approach - We designed the platform around the unit MLtwist actually sells: a customer project. Each project carries its data sources, labeling instructions, workforce assignments, quality rules and delivery format, so onboarding a new customer is configuration rather than a new codebase. - Processing runs as containerized workloads on Kubernetes. When a project needs data pulled, transformed, labeled or packaged, the platform triggers the right custom pipeline as jobs that scale up for large batches and back down when idle. - Around that core we built workforce management for labelers and reviewers, labeling project setup and assignment, automated quality control, and versioned, tracked data in a secure environment. ### Scope - Project and customer management - Workforce management for labelers and reviewers - Labeling project setup, assignment and QC - Custom pipelines triggered per project on Kubernetes - Containerized processing workloads that scale per job - Versioned, tracked data in a secure environment Stack: Kubernetes, Containers, Data pipelines, Human-in-the-loop labeling, Postgres, Object storage Outcome: The platform is the operating system for MLtwist’s delivery — the foundation for its public and private sector data work. ## Multimodal AI data for the U.S. TSA URL: https://axiointelligence.com/work/tsa-multimodal-ai-data/ Client: MLtwist · for the U.S. Transportation Security Administration · Sector: Government · aviation security · Status: Contract awarded Feb 2026 Video, audio, text, images and 3D DICOS scans — processed and labeled for threat-detection models, on the platform we built. Our role: MLtwist holds the TSA contract. Axio engineered the MLtwist platform the work runs on. Challenge: The TSA needed multimodal data — video, audio, text, images and 3D scans in the DICOS security-imaging standard — preprocessed, labeled and packaged to develop threat-detection models for aviation screening. ### Approach - The MLtwist platform we built already treated every customer as a configurable project, so the TSA work was set up as project configuration: DICOS and multimodal inputs, labeling instructions, reviewers and delivery formats. - Kubernetes-triggered pipelines handled preprocessing and transformation for each data type, feeding human-in-the-loop labeling with automated quality control, and packaged results as JSON or DICOS. - Every dataset stayed versioned and tracked in a secure environment, from raw input to delivered package. ### Scope - Multimodal ingest: video, audio, text, image and 3D DICOS - Preprocessing and transformation pipelines - Human-in-the-loop labeling with automated QC - Versioned, tracked data in a secure environment - Per-project setup of inputs, instructions, reviewers and formats - Delivery as JSON or DICOS packages Stack: Kubernetes, Multimodal pipelines, DICOS, Human-in-the-loop labeling Outcome: MLtwist reports that its pilot cut processing time from eight weeks to three. In February 2026 the TSA awarded MLtwist a $590K contract for multimodal AI data labeling and processing. Sources: - MLtwist: TSA awards $590K contract (Feb 2026): https://mltwist.com/2026/02/03/mltwist-awarded-590k-tsa-contrat-for-multi-modal-labeling/ ## PRIMED: fast materials-data search for U.S. Department of Energy researchers URL: https://axiointelligence.com/work/doe-materials-data-search/ Client: MLtwist · for the U.S. Department of Energy (SBIR) · Sector: Government · materials science · Status: Delivered PRIMED, a search and indexing platform that lets scientists share materials data and find materials by recorded and projected properties — in seconds. Our role: Axio led an augmented engineering team that built PRIMED for MLtwist. Challenge: Materials scientists generate measurements across labs and instruments, while computational models project properties for materials nobody has measured yet. Finding a candidate material meant digging through papers, spreadsheets and siloed datasets — with no way to search what has been measured and what has been predicted together. ### Approach - We built PRIMED (Preparation & Requisition of Integrated Microscopy + Energy Data): a shared platform where scientists publish their materials data, findings and measurements, with an indexing and search layer fast enough to explore interactively. - Every material is indexed on both recorded measurements and projected, model-predicted properties, so a single query spans both. A scientist looking for a material that refracts specific wavelengths of light and also withstands a given temperature can ask for exactly that. - Multi-property, range-based queries across optical, thermal and other properties return ranked candidates in seconds. Axio led an augmented engineering team working alongside MLtwist to design and ship it. ### Scope - Shared repository for materials data, findings and measurements - High-speed indexing of recorded and projected properties - Multi-property, range-based search (e.g. optical + thermal) - Ranked candidate results for interactive exploration - Publishing workflow so scientists can share new datasets - Augmented engineering team working with MLtwist Stack: Search & indexing, Scientific data models, Data pipelines, Web application Outcome: Researchers can find candidate materials by the properties they need — measured or projected — instead of by the paper they happened to read. PRIMED was funded through MLtwist’s $198,834 U.S. Department of Energy SBIR award. Sources: - MLtwist: DOE SBIR award for PRIMED (Jan 2022): https://mltwist.com/2022/01/20/mltwist-wins-doe-deal/ ## Dealer telematics and lot management, rebuilt from scratch URL: https://axiointelligence.com/work/dealer-telematics-platform/ Client: Automotive telematics provider · confidential · Sector: Automotive · dealer telematics · Status: In production A new telematics platform for car dealerships — built from zero, on lower-cost hardware, field tested and in production. Our role: Axio built the platform end to end — software, hardware sourcing, field testing and production rollout. Challenge: A telematics provider serving car dealerships needed to replace its existing telematics and lot-management platform. Dealers depend on it to know where every vehicle on the lot is, whether one is moving when it shouldn’t be, and whether devices are healthy — and the incumbent hardware was expensive. ### Approach - We built the entire platform from scratch: device firmware and cloud ingestion, the event and alerting engine, and the dealer-facing lot-management application. - We sourced hardware manufacturing and devices below the client’s existing unit cost, then field tested on real dealer lots before rolling out to production. - Devices send health checks and movement events over a global cellular IoT connectivity platform. The platform turns those events into geofence-breach and speeding alerts, and keeps a live view of every vehicle on the lot. ### Scope - Platform built from scratch: firmware, cloud and dealer app - Hardware sourcing and manufacturing below prior unit cost - Field testing on dealer lots, then production rollout - Device health checks and movement events over cellular IoT - Geofence-breach and speeding alerts - Lot management: live vehicle location and inventory Stack: Cellular IoT, Embedded firmware, Device-to-cloud, Event & alerting engine, Geofencing, Web application Outcome: The new platform is in production, replacing the client’s previous telematics and lot-management system on lower-cost hardware. ## An AI-first platform rebuild behind a company’s pivot URL: https://axiointelligence.com/work/ai-first-platform-pivot/ Client: Connected-device software company · confidential · Sector: Connected-device software · Status: In production · cloud and on-prem One platform for a new embedded Linux product, running on a laptop, in AWS and air-gapped on-prem. Built so the client’s own team could finish the overhaul with AI in under two weeks. Our role: Axio led the client’s software team through the pivot: we scaffolded the infrastructure and platform, then handed the team a codebase built for AI-assisted development. Challenge: After several years building a device-management platform, a connected-device software company decided on a major pivot: a new embedded Linux operating system for device makers. The platform behind it had to change too, and the team needed a clear path from the old platform to the new one. It also had to run in three places: on each developer’s laptop, in AWS, and inside customers’ own air-gapped facilities with no connection to the cloud. ### Approach - We designed one platform that runs the same way everywhere. Developers get a local copy that matches the AWS cloud environment, so what works on a laptop works in production. On-prem customers get an edition that runs on Kubernetes (K8s or K3s) inside air-gapped networks, managed without reaching the internet. - We scaffolded the initial infrastructure and platform as code, with reproducible environments and a one-command local setup, so every environment is built from the same definitions. - Then we put the guardrails in place that make AI-assisted development safe: agent instruction files that encode the project’s conventions, architecture docs and task specs that the team and its AI agents work from, and test suites and CI gates that block any change that doesn’t pass. - Finally, we walked the team through the new project and how to work in it AI-first, so they had a running start rather than a blank repository. ### Scope - Infrastructure and platform scaffolding as code - Local developer environment matching AWS - Air-gapped on-prem edition on K8s and K3s - Agent instruction files and project conventions - Test suites and CI gates as guardrails - Architecture docs, task specs and team onboarding Stack: AWS, Kubernetes, K3s, Infrastructure as code, CI/CD, AI coding agents Outcome: With the guardrails in place, the client’s own team finished the platform overhaul in under two weeks. The new platform runs in production on AWS and air-gapped on customer premises, and every developer runs the same stack locally. # Blog ## Same-SIM satellite IoT: what works today URL: https://axiointelligence.com/blog/satellite_iot_same_sim/ Published: 2026-10-01 · Author: Brice Ayres · Tags: IoT, satellite, NTN, connectivity Short answer: Satellite IoT can now run on the same SIM as cellular through 3GPP Release 17 NTN, but only when the module, the connectivity provider's satellite agreement, the antenna and the firmware all support it. In our testing, many providers advertising satellite IoT connected intermittently or not at all, so ask for a test SIM and written coverage before you commit. Cell networks were built where people are. A lot of the things worth monitoring aren't: wellheads, pump stations, livestock water tanks, rail cars, shipping containers, sensors on a ridgeline. When a device drives or floats out of coverage, it goes quiet, and you find out what happened later, if at all. Until recently, the answer was a second radio. You added a satellite modem next to the cellular one, signed a second contract, paid for a second data plan, and built a second pipeline into your cloud. It worked, but it doubled the hardware, the firmware paths and the billing for a link most devices only need occasionally. That's changing. Satellite networks can now act like another cell network, reachable from a standard IoT module on the SIM already in the device. That's the promise, anyway. Here's what makes it hard, and what we've learned from testing it. ### Why is satellite IoT harder than cellular? Satellite isn't cellular with a longer range. Your device and your software need to handle it differently. - **Small payloads.** Think tens to a few hundred bytes per message: a position, a tank level, an alarm. It isn't a pipe for logs, photos or firmware updates. - **Latency in seconds, not milliseconds.** Messages can take seconds or longer to land, and some networks queue them until a satellite is overhead. Anything that expects a fast round trip, like a chatty protocol or a TLS handshake on every message, needs to be rethought. - **The device has to know where it is.** Under the 3GPP standard for satellite IoT, the device uses a GNSS position fix to correct for satellite motion and timing before it transmits. Many modules can't run GNSS and the cellular radio at the same time, so firmware has to sequence them, and that costs time and battery. - **Sky view and antennas.** The device needs a clear view of the sky, and satellite runs on different frequency bands than terrestrial LTE. An antenna tuned for cellular, mounted inside a steel enclosure, won't close the link. - **Power.** Transmitting to a satellite takes more energy per message than transmitting to a tower a few miles away. Battery-powered devices need a budget for how often they're allowed to use it. - **Cost per byte.** Satellite data costs far more than cellular. Without rules about which messages may go over satellite, a misbehaving device can burn through a month's budget in an afternoon. - **Failover is your problem.** Something has to notice that cellular is gone, decide satellite is worth the cost, switch, queue what can wait, and switch back. Most of the time that's your firmware, not the network. None of these is a deal-breaker. All of them have to be designed for from the start, not added after the device ships. ### How does same-SIM satellite IoT work? The standards work that made this possible is 3GPP Release 17, which defined non-terrestrial networks (NTN) for NB-IoT and LTE-M. In plain terms, a satellite network can now present itself to a device the way a cell network does. From there, it works much like roaming. If your connectivity provider has an agreement with a satellite network operator, the satellite network shows up as one more network your SIM is allowed to join. The device keeps its SIM and its identity, and traffic can come back to your cloud through the same path as your cellular traffic. You keep one SIM, one device record and one pipeline. The conditions matter, though. You need: 1. **A module that supports NTN**, with firmware that actually has it enabled, on the right satellite bands. 2. **A connectivity provider with a working satellite agreement** that covers the places your devices go. 3. **The right antenna and mounting** for the satellite bands, with a view of the sky. 4. **Firmware that knows the difference** between a cellular link and a satellite one, and behaves accordingly. When all four line up, it's a real improvement over a second radio. When one is missing, you get a device that says "satellite-ready" on the box and never sends a byte over satellite. ### "We have satellite" can mean very different things Over the past several months we've tested satellite offerings from several connectivity providers on real devices, on the bench and in the field. Many advertise satellite IoT. In our testing, a lot of them connected only intermittently, in limited areas, or not at all. The marketing was ahead of the network. That isn't always bad faith. This is a new technology, and an announced partnership, a pilot and a commercial service are three different things that tend to share one press release. But if you're planning a product around it, the difference matters. Before you commit, ask: - **Is it NTN on the same SIM, or a separate satellite modem and contract?** Both are valid. They're very different engineering and cost decisions. - **Is it commercially live where your devices operate today?** Not "launching," not "in trials." Ask for the coverage area in writing. - **Which modules and firmware versions are approved to use it?** Can you buy those modules in volume now? - **Can we test it before we sign?** A provider confident in its network will put a test SIM in your device. If they can't, that tells you something. - **Who handles failover?** Does anything in their stack help, or does "support" just mean the SIM is allowed to try? - **What does a message actually cost,** and does satellite data come back through the same endpoint as cellular, or a separate one you have to integrate? If the answers are vague, plan as if satellite isn't there yet. ### What we're building After a lot of testing, we found one network that connects consistently. We're working toward a partnership with them. On top of that partnership, we plan to build a platform that lets anyone get satellite connectivity for their devices without becoming a satellite expert first. The idea is simple: pick a device that supports satellite, activate it, and have its data arrive in your cloud, with the cellular-to-satellite switching and the message budgeting handled for you. No second radio, no separate pipeline, no months of integration work to find out whether it connects. It's early. There's no name or launch date yet, and we'll share more as the partnership and the platform take shape. If you have devices that need to work beyond cell coverage and want to be among the first to try it, email us at [contact@axiointelligence.com](mailto:contact@axiointelligence.com). We'd like to hear what you're connecting. If you need cellular-to-satellite failover designed into a device now, that's [work we already do](https://axiointelligence.com/services/cellular-satellite-connectivity/). ## What actually sets your z/OS software bill URL: https://axiointelligence.com/blog/zos_peak_software_cost/ Published: 2026-10-01 · Author: Brice Ayres · Tags: mainframe, z/OS, COBOL, cost reduction Short answer: IBM's monthly-licensed z/OS software, including z/OS, CICS, Db2, IMS and MQ, is usually billed on the highest rolling four-hour average of MSU use in the month, so the busiest four hours set the bill. To lower it, find the batch jobs in that window with SCRT, SMF and scheduler data, rank them by share of the peak rather than total CPU, then reschedule them or rebuild them in the cloud one at a time with a dual run. Most companies still on an IBM mainframe know the monthly software bill is large. Fewer know which four hours of the month set it. That matters because the bill for IBM's monthly-licensed software on z/OS isn't based on how much you use the machine over the month. It's based on the busiest stretch. If you can find what runs in that stretch and move it, the bill comes down. You don't have to rewrite everything at once to get there. This post covers how the peak works, how to find the jobs inside it, and the two levers for lowering it. ### The bill is set by four hours, not thirty days IBM's Monthly License Charge (MLC) software, which includes z/OS itself, CICS, Db2, IMS and MQ, is usually billed on **sub-capacity pricing**. Capacity is measured in MSUs (millions of service units, IBM's unit of processing capacity). For each product, IBM looks at the LPARs where it runs and takes the **highest rolling four-hour average** of MSU use during the month. That one number sets the charge. The data comes from your own system. SMF records track CPU use continuously, and IBM's Sub-Capacity Reporting Tool (SCRT) turns them into the monthly report you submit. So the peak is no secret. It's in a report your systems programmers already produce. Two consequences follow: - **Quiet hours are cheap.** A job that burns a lot of CPU at 3 a.m. on a Sunday, when nothing else is running, may not add a dollar to the bill. - **One busy window is expensive.** A job that adds modest CPU at the moment of the monthly peak raises the bill for the whole month. The rolling four-hour peak is common, but it isn't the only model. Depending on the vendor and the contract, mainframe software can be priced on the peak, on the machine's full capacity, on total consumption, or through an enterprise agreement, and third-party software from Broadcom and BMC adds its own terms. The details vary too much to generalize. What holds under every model is simpler: the bill follows the work that runs on the mainframe. Less work means less to license and less to run. ### Which jobs are in the peak? For most manufacturers, distributors and retailers, the monthly peak isn't the online day. It's batch: - Nightly inventory and replenishment runs - Pricing and promotion feeds to stores, e-commerce and partners - Month-end close and the reports that go with it - Extracts to the data warehouse, planning tools and EDI partners Many of these jobs follow the same pattern: **read IBM data, write a file, feed or report.** They don't post transactions or change the system of record. They're islands, and islands are the easiest thing to move. To see which ones are in your peak, line up three sources: 1. **SCRT reports** for the last several months, to find the peak hour each month and which LPARs drive it. 2. **SMF records** for those windows: job-level accounting (type 30) and workload activity (type 72) show which jobs and service classes were consuming CPU at the time. 3. **The scheduler export** (Control-M, CA-7, IBM Z Workload Scheduler or similar), to see what each job depends on, what depends on it and when it's allowed to run. ### Rank by share of the peak, not total CPU This is where most cost exercises go wrong. The instinct is to go after the biggest CPU consumers. But the bill doesn't care about total CPU. It cares about CPU in the peak window. So rank every job on two axes: - **Dollars**: how much of the monthly peak this job accounts for, and what that's worth at your MLC and third-party rates. - **Risk**: what breaks downstream if the job is late or wrong, and how well anyone still understands it. The best first candidates sit high on dollars and low on risk: big contributors to the peak with simple inputs, simple outputs and a clear owner. ### First lever: move the job in time If your contract is priced on the peak, the cheapest fix is often a scheduling change. If a job is in the peak only because it was always scheduled at that time, and nothing downstream needs it then, moving it to a quiet window can lower the peak without new code. Check this before anything else. Also check whether your team already caps capacity (defined capacity or group capacity limits). Capping lowers the bill by slowing work down when the four-hour average hits the limit, which is exactly what you don't want for time-sensitive batch. Some work can also run on **zIIP** specialty engines, which don't count toward the MSU figure for software pricing. Db2 distributed requests, Java and some utilities qualify. Ordinary COBOL batch generally doesn't. ### Second lever: move the job off the box When a job can't move in time because the stores, the warehouse or a partner needs its output by a deadline, rebuild that one job in the cloud: 1. **Replicate only what it reads.** Land the job's input files or tables in AWS, Azure or Google Cloud, as a scheduled extract or a change data capture feed. Time the replication outside the peak, or you've just moved CPU from one job to another. 2. **Rebuild the logic** against Postgres or the database you already run, with the same output format the downstream systems expect. 3. **Dual-run.** Run the new job alongside the old one every cycle and compare the outputs until they match, including month-end and other edge cases. 4. **Turn the old job off** once the outputs match and the people downstream have signed off. Then the next job. Everything else keeps running on the mainframe while you do this. You're taking load off the box one job at a time, with evidence at every step, and each job you move makes a full migration smaller if and when you choose one. ### Start around the core Start with the jobs around the core, not the core itself. Programs that post transactions are bigger, riskier moves, best made once the team and the cloud platform have a track record from the easier jobs. And some systems, like a bank's deposit core or a government benefits system, aren't candidates for this approach at all. ### Does this apply to IBM i? IBM i (the old AS/400) isn't licensed this way, and there's no rolling four-hour peak. The cost there is licensing and hardware, and usually people: decades of RPG and COBOL batch that fewer people understand every year. The method (inventory the jobs, pick one, dual-run, turn it off) is the same. The ranking is different: by who screams if it's late, and by whether the person who understands it is still around. ### How we run it We start with a diagnostic: your contracts and invoices, the scheduler export and usage reports on z/OS, or the job scheduler and programs on IBM i. The result is a ranked job list, a migration plan and a baseline of today's licensing and compute spend. There's no upfront fee; we're paid a percentage of the net savings against that baseline, after the cost of running the work in the cloud. More on the offer: [mainframe & IBM i migration](https://axiointelligence.com/services/mainframe-migration/). If you're working through the same question for your cloud bill, the [cloud cost diagnostic](https://axiointelligence.com/services/cloud-cost-reduction/) is the same idea pointed at a different invoice. ## How we usually find 30–50% of a cloud bill in the first week URL: https://axiointelligence.com/blog/quietly_cutting_cloud_bills/ Published: 2026-05-12 · Author: Brice Ayres · Tags: FinOps, AWS, cost reduction, Kubernetes Short answer: Most cloud waste hides in orphaned resources, forgotten environments and over-provisioned workloads, not where leadership expects. In the first week we inventory every billable resource (typically 5–15% of the bill is things that should have been turned off), right-size the top spenders from 30 days of p95 data, and only then decide on savings plans and spot. The result is a cut-now, cut-next and guardrails list delivered as pull requests. Every cloud cost engagement starts the same way: finance is alarmed, engineering is defensive, and nobody has a complete picture of what's actually running. This isn't because engineers are sloppy. It's because cloud spend grows in the seams between teams, between environments, and between the things people built and the things they forgot to turn off. Almost every time, the waste is somewhere different from where leadership thinks it is. Here's the rough playbook we run in the first week of a typical cost engagement, and what we usually find. ### Day 1–2: where is the waste hiding? The first thing we do is build an inventory of every billable resource, tagged by service, environment, owner, and last-touched. Most teams don't have this. They have a billing dashboard, which is not the same thing. What we're looking for: - **Orphans.** RDS replicas attached to a primary that was migrated two years ago. ELBs pointing at deregistered targets. EBS volumes not attached to any instance. NAT gateways nobody can explain. - **Untagged sprawl.** Resources without owners. If nobody owns it, nobody will turn it off. We start here because it's almost always cheap to delete. - **Forgotten environments.** A staging cluster that's been at production scale since a load test in 2023. A dev account someone spun up for a POC three roles ago. This inventory work isn't glamorous. It is reliably 5–15% of the bill, just from things that should have been turned off and weren't. ### Day 3–4: the workloads that are wrong-sized This is where engineering teams usually get tense, because it feels like a critique. It's not. Workloads are almost always over-provisioned because the cost of being wrong in the other direction is a 3am page. We look at p95 CPU, p95 memory, and request patterns over 30 days for the top spenders. We're trying to answer: - **Are you paying for the worst case all the time?** Most workloads have a 2–4x spread between average and peak. Autoscaling — real autoscaling, not "we have an ASG configured" — closes most of this. - **Is your compute the right shape?** Memory-optimized instances running CPU-bound workloads. GPU instances doing CPU work. ARM workloads on x86 because Graviton wasn't a thing when the instance type was picked. - **Is your storage tier right?** EBS gp2 vs gp3 alone is usually a 20% storage savings. S3 standard for cold archival data. Logs in CloudWatch instead of S3 with a lifecycle policy. This is also where we usually find Kubernetes clusters running at 18% utilization with no horizontal pod autoscaler, no cluster autoscaler, and node groups sized for a peak that happens 4 hours a week. ### Day 5: the commitment + spot conversation Once we know what your actual baseline is, we can talk about reserved instances, savings plans, and spot. This conversation in the wrong order is what gets people stuck on 3-year commits for the wrong instance families. Rough heuristic: anything that runs 24/7 should be on a savings plan. Anything stateless and tolerant of interruption should be on spot. Anything with predictable bursts can usually move to scheduled scaling. For Kubernetes specifically: Karpenter on spot, with a small on-demand base, will typically take a node bill down 50–70% with almost no operational change. ### What we ship at end of week one A short document with three sections: 1. **Cut now.** Things we'd delete or resize today. Usually 15–25% of the bill, often more. 2. **Cut next.** Architectural changes that take a few weeks but compound. Spot migration, Graviton, storage tiering, data egress. 3. **Stop the bleed.** Guardrails so the waste doesn't grow back. Budget alerts per team, anomaly detection, tag enforcement, a quarterly review cadence. The handoff isn't "here's a slide deck." It's PRs against your Terraform, a tagging policy, a budget dashboard, and a list of "next" items prioritized by dollars saved per engineering hour. ### What we don't do We don't run optimization tools that promise to do this in the background. They're a great way to get a 5% improvement and feel like you're done. The actual savings — the 30–50% — comes from looking at workloads as an engineer who understands what they do, not as a script looking at metrics. If your cloud bill has quietly become someone's full-time problem, this is [the kind of thing we do](https://axiointelligence.com/services/cloud-cost-reduction/). ## Migrating clouds without the migration freeze URL: https://axiointelligence.com/blog/migration_without_downtime/ Published: 2026-04-18 · Author: Brice Ayres · Tags: migration, AWS, GCP, Postgres, infrastructure Short answer: Cloud migrations finish on time when there is no cutover weekend. Map everything you run first, migrate one small slice to prove the supporting pieces, move stateful services through dual writes and shadow reads, shift traffic gradually at the load balancer, and turn the old system off two weeks after it stops taking traffic. For a company with about a dozen services, plan on 8–14 weeks while the team keeps shipping features. The reason most cloud migrations slip is that they're scoped as an event. A weekend. A holiday. A window where the team can take a deep breath, flip the switch, and pray. We don't run them that way anymore. After enough migrations, the pattern is clear: the projects that finish on time are the ones where there is no cutover weekend at all. The new system runs alongside the old one for weeks, taking progressively more traffic, until the old system has nothing left to do and you turn it off on a Tuesday afternoon. Here's how we structure it. ### Step 1: forget about the destination for a minute Before we talk about which cloud to move to, we map what you actually have. This sounds obvious. It isn't. Most teams have a partial picture: the services they ship are well-understood, but the queues, jobs, dashboards, and integrations that grew up around them are not. We've inherited migrations that were 80% done when someone discovered a Lambda that nobody owned was pushing critical reports to a third-party SFTP server. The deliverable for this phase is a single diagram and a CSV. The diagram shows every running thing and what it talks to. The CSV lists every billable resource and its owner. If nobody owns it, that's a finding. ### Step 2: pick the smallest viable first slice The most common failure mode is "let's lift-and-shift everything, then optimize later." This works on whiteboards. It fails in practice because you're moving everything at once and you have to coordinate everyone at once. Instead, we pick a slice. Usually it's: - One stateless service - Its database (or a logical subset) - The CI pipeline that ships it - The monitoring + on-call wiring The slice has to be small enough to ship in 2–3 weeks, and complete enough that you can actually run it in production on the new cloud. The point isn't to migrate this service — it's to prove out every supporting piece (IaC, CI, secrets, observability, network) on something low-risk. ### Step 3: dual-write, dual-read, then cut For services with state, we almost always go through a dual-write phase: 1. **Dual-write.** New writes go to both old and new databases. Reads still come from the old. We run this until backfill of historical data is complete and the two systems are byte-for-byte consistent. 2. **Shadow read.** Reads start hitting both systems. We compare results, log discrepancies, and don't return the new system's data to users. This catches every subtle behavior difference — character encoding, sort order, NULL handling — without anyone noticing. 3. **Flip reads.** Move the read source. The old database is now a hot backup. 4. **Stop writing to the old.** Now you can decommission on your own time. For Postgres specifically: logical replication via `pglogical` or AWS DMS, with row-level consistency checks. For Mongo, change streams. For event-sourced systems, replay from your event log and dual-publish. The point is that at every step, you can roll back to the previous state in seconds, not hours. ### Step 4: traffic shifting at the edge Once a service has its new home, we shift traffic gradually at the load balancer or DNS layer. Usually: - 1% for an hour - 10% for a day - 50% for a few days - 100% If anything goes sideways, the rollback is moving the weight back. There is no "let's revert the deploy and restore from backup." This is also where we catch the long-tail issues you couldn't catch in staging: cold caches, region-specific latency, IAM policies that work for the test account but not the production one. ### Step 5: decommission, then prove it A migration isn't done when the new system is taking 100% of traffic. It's done when the old one is off, the bill is gone, and the IaC for it is deleted from the repo. We typically wait two weeks after 100% before we tear things down. After that, the AWS account gets a `terraform destroy`, the IAM users get revoked, and the DNS records get pruned. The PR to remove the old code is sometimes the most satisfying one in the whole project. ### How long does a cloud migration take? A reasonable benchmark for a single-product company with a moderately complex backend (a dozen services, two databases, a queue or two, a frontend, and CI/CD): **8–14 weeks**, with the team continuing to ship features the whole time. What blows that estimate is almost always the same thing: undocumented dependencies discovered halfway through. The inventory step is the cheapest insurance you can buy. If you're staring at a migration plan that has a single weekend cutover on it, we'd recommend re-scoping it. There's no reason to take that risk. Here's [how we run migrations](https://axiointelligence.com/services/legacy-modernization/). ## Infrastructure-as-code that survives the team that wrote it URL: https://axiointelligence.com/blog/iac_that_survives_the_team/ Published: 2026-03-22 · Author: Brice Ayres · Tags: IaC, Terraform, OpenTofu, platform engineering Short answer: Terraform and OpenTofu codebases rot because of their architecture, not the tool. Split state by blast radius, write modules only for real repeated patterns, keep state in a locked remote backend with no applies from laptops, post a plan on every pull request, run drift detection, and limit policy-as-code to a few guardrails. The test is whether a new senior engineer can ship a safe change in their first week without asking anyone. We've inherited a lot of Terraform. The pattern is depressingly consistent: a repo started by a founder or an early platform engineer who knew exactly what they meant, then accreted layers as the team grew, then became something nobody wanted to touch. It's never the tool's fault. It's the architecture. Here's what we've found makes the difference between IaC that ages well and IaC that becomes a museum. ### The single biggest mistake: one-state-to-rule-them-all A surprising number of teams keep everything — networking, databases, IAM, apps, secrets — in a single Terraform state. It works for a while. Then: - A plan takes 4 minutes - An apply locks the state for everyone - One person's change to a Lambda blocks the database team - A `terraform destroy` becomes impossible to reason about The fix is boring: **split state by blast radius**. Things that change together stay together; things that don't, don't. A typical layout we use: ``` stacks/ platform/ # VPC, subnets, IAM baseline, KMS, route53 data/ # RDS, Redis, S3 buckets, replication shared-services/ # observability, secrets, service mesh apps/ api/ worker/ web/ ``` Each stack has its own state. Cross-stack dependencies go through `terraform_remote_state` or, better, a small set of well-known SSM parameters or Secrets Manager entries. Apps don't know how the VPC is built; they just ask for the subnet IDs. This one change — splitting state — typically cuts plan times by 10x and lets multiple engineers work in parallel without stepping on each other. ### Modules, not abstractions The second pattern that hurts is the urge to wrap everything in a custom module. "We don't want people writing raw `aws_instance` resources." So now there's a homegrown `company_ec2_instance` module that's a thin wrapper around the AWS one — with three opinions baked in and twelve variables that nobody understands. A year later, somebody needs to do something the module doesn't support. They open the module, get confused, copy-paste-modify it inline, and now you have three slightly different versions of the same thing. Our heuristic: write modules where you have a real, repeated pattern that benefits from being named — a microservice, a Postgres cluster with the team's standard backup config, an EKS node group with the right tags. Don't write modules to "wrap" upstream resources. The upstream provider is fine. ### Where should Terraform state live? If your team has more than two engineers, your state belongs in a remote backend with locking. S3 + DynamoDB works fine. Terraform Cloud, Spacelift, or Env0 work better — they give you per-PR plan visibility and a permission model. Whatever you pick: nobody applies from their laptop. Period. The mental shift is: the IaC repo isn't a script you run, it's the source of truth that an applier acts on. PRs run plans. Merges run applies. State doesn't live on anyone's machine. ### The pre-merge `plan` is your code review Every PR should post a plan output as a comment. Atlantis, Terraform Cloud, and Spacelift all do this out of the box. Once you have it, your reviewers can actually see what's about to happen — not just what the HCL says, but what AWS is about to do. This single feedback loop catches more would-be incidents than any policy-as-code framework. It's also how junior engineers learn what their changes mean. ### Drift detection isn't optional Manual changes in the console will happen. Somebody will be on-call at 2am, fix a thing, and forget to bring it back into code. Without drift detection, your IaC is silently lying to you, and the next person to plan in that area is going to get a surprise. Cheap version: a nightly scheduled job that runs `terraform plan` and posts to Slack if anything's drifted. Better version: drift detection in Terraform Cloud / Spacelift with auto-remediation policies. ### Policy-as-code, applied lightly OPA / Sentinel are great. They're also great at making your platform team into the bottleneck if you over-rule them. The policies that have paid for themselves in our work: - No public S3 buckets without an explicit `:public` tag - No security groups with `0.0.0.0/0` on non-HTTPS ports - No resources without an `owner` tag - No RDS without backups enabled That's it. Three or four guardrails that catch real mistakes. Save the rest for after the team is mature enough to want them. ### Handoff is part of the work When we finish an IaC rewrite, we don't just hand over a repo. We hand over: - A README that explains the layout in two pages - A "how to do common things" doc — add a service, add a queue, rotate a secret, bootstrap a new env - An onboarding PR that a new engineer can use to learn by doing - A decision log (`docs/decisions/`) explaining why things are the way they are The test for a good IaC codebase is: can a new senior engineer ship a safe change in their first week, without asking anyone? If yes, you're done. If no, the work isn't finished. It's the bar we hold our own [infrastructure-as-code work](https://axiointelligence.com/services/cloud-infrastructure-iac/) to. ## Kubernetes, the boring way URL: https://axiointelligence.com/blog/kubernetes_the_boring_way/ Published: 2026-02-14 · Author: Brice Ayres · Tags: Kubernetes, EKS, GitOps, platform engineering Short answer: A production Kubernetes cluster that a team of eight can run without a dedicated platform engineer needs a managed control plane (EKS, GKE or AKS), one cluster per environment and region, Terraform or OpenTofu for the bootstrap, and GitOps with ArgoCD or Flux for every change after that. Keep secrets in a real secret manager through External Secrets Operator, give every service metrics, logs and traces by default, and upgrade one minor version per quarter. Kubernetes is not a problem. Kubernetes set up by someone who's no longer on the team is a problem. Most of the k8s clusters we inherit have the same two flaws: they were assembled from a hundred good blog posts, and the person who assembled them is gone. The cluster runs, mostly. Upgrades terrify everyone. Nobody can quite explain why ingress works the way it does. Here's the stack we default to when we bootstrap a cluster, and the rules of thumb that come with it. The goal is a cluster a team of 8 can run without a dedicated platform engineer. ### The base - **Managed control plane.** EKS, GKE, or AKS. Self-managed control planes are a discipline. Most teams don't need that discipline. - **One cluster per environment, per region.** Not one cluster with namespaces for prod/staging/dev. The blast radius and IAM boundaries are too important. - **Node pools split by workload.** A general pool, a memory-intensive pool, a spot pool for stateless workloads. Use Karpenter (EKS) or the native cluster autoscaler. Don't pre-size for peak. ### Bootstrap with IaC, then never touch it manually The cluster, node groups, IAM, networking, KMS keys — all in Terraform/OpenTofu. The in-cluster bits — controllers, ingress, monitoring agents — installed via a bootstrap pipeline that the IaC kicks off. After bootstrap, the only thing that changes the cluster is GitOps. `kubectl apply` from a laptop is forbidden. If you can't do it through a PR, you don't do it. ### GitOps with ArgoCD or Flux We use ArgoCD more often because the UI is helpful when you're explaining the system to a new team member. Flux is great if you prefer everything to be CRD-driven and don't need the dashboard. The structure we like: ``` gitops/ apps/ # one folder per app, with kustomize overlays per env api/ base/ overlays/ staging/ prod/ platform/ # cluster-wide things cert-manager/ external-dns/ ingress-nginx/ kube-prometheus/ ``` ArgoCD watches the `apps/` and `platform/` paths. Engineers ship by opening a PR against the overlays. Promotions between environments are PRs, not buttons. ### Ingress: pick one, stop arguing `ingress-nginx` is the boring choice and the right choice for most teams. AWS Load Balancer Controller if you specifically want ALBs per ingress. Don't run three ingress controllers because three different teams had opinions on the same week. If you genuinely need a service mesh — and most teams don't — Linkerd. It has the lowest operational tax of any mesh we've run. ### Where should Kubernetes secrets live? Don't use Kubernetes Secrets as your secret store. Use a real secret manager (AWS Secrets Manager, GCP Secret Manager, or Vault) and pull into the cluster via External Secrets Operator. This way: - Rotation happens in one place - Audit logs are real - Compromising kube doesn't compromise everything ### Observability The default we ship: - **Metrics:** Prometheus + Grafana, deployed via `kube-prometheus-stack`. One Grafana per cluster, federated up if you have many. - **Logs:** Loki or whatever your team is already paying for (Datadog, Splunk, etc). Don't run your own Elasticsearch unless logs are your business. - **Traces:** OpenTelemetry collector → your tracing backend of choice. The point isn't the tools. The point is that on day one, every new service automatically gets metrics, logs, and traces without the team having to opt in. ### How often should you upgrade Kubernetes? The single most important question we ask a team that has Kubernetes: "what's your upgrade story?" If the answer is "we'll figure it out when we have to" — they're going to have a bad time. Kubernetes ships a new minor version every four months and EOL'd versions get scary fast. Our default: 1. **Cadence:** upgrade one minor version per quarter. 2. **Process:** dev → staging → prod, one week apart each, with a checklist. 3. **Documentation:** every upgrade gets a one-page runbook of what changed and what broke. The first upgrade is painful. The third one is boring. That's the goal. ### Hardening: enough to sleep at night - Pod Security Standards (`restricted` for app namespaces) - Network policies that default-deny ingress between namespaces - IRSA / Workload Identity instead of node-level IAM - Audit logging shipped off-cluster - A schedule for re-running CIS benchmarks You don't have to be CKS-certified to run a safe cluster. You do have to do these five things. ### What this gets you A cluster a small team can run for a year without major incidents. Upgrades that take an afternoon instead of a sprint. A platform that new engineers can understand from a README, not from interviewing the founder. If your Kubernetes setup currently requires the founder to operate it, that's the kind of [cleanup we do](https://axiointelligence.com/services/devops-sre/).