Skip to content

DevOps Engineer · Platform Engineer · Toulouse, France

I build platforms that keep running when I’m not there.

Four years of cloud infrastructure on assignments in aerospace and industry: AWS, Azure, Kubernetes, secure CI/CD, costs under control. And a personal platform in production, operated by AI agents. This site is part of it.

Boubacar Soumaré
in production · 24/7 · self-hosted infrastructure
  • Internal platform

    A self-hosted platform, in production, whose users are AI agents: written rules, generated state, test benches, alerts. The same constraints as a team of developers.

    see the full proof

  • Cloud and infrastructure as code

    On assignment: AWS costs cut by 65 to 75 % without a single incident, isolated Azure environments from day one. On the platform: DNS, TLS and firewall described as code, nothing applied without a reviewed plan.

    see the full proof

  • Kubernetes and GitOps

    On assignment: EKS on a secure data platform. On the platform: manifests in the repository, reconciled by a GitOps controller, pinned versions; written, bench-tested.

    see the full proof

  • CI/CD and supply-chain security

    On assignment: static analysis, SonarQube and end-to-end tests in Jenkins pipelines. On the platform: secret scanning, dependency audit, an image built then scanned on every push.

    see the full proof

  • Observability and SLOs

    An error workflow wired to every production workflow, a probe each morning comparing declared state with real state. Quantified SLOs are the next step.

    see the full proof

  • AI agents

    Agents that act without me through MCP and agent SDKs, and autonomous workflows in production, each covered by a check.

    see the full proof

  • Governance

    An entry point every agent reads first, a generated system state, test benches, a work registry: I can verify what was done.

    see the full proof

Delivery

On assignment and on my platform, the same standard

Measured results for industrial clients, and the delivery chain of my platform as it stands today.

On assignment

A European aircraft manufacturer · AWS

AWS costs cut by 65 to 75 %

0 incidents

Audit of real provisioning, service by service, then a reduction in two one-week steps validated by the alarms. No production interruption.

A European aircraft manufacturer · AWS

Securing a multi-service platform

EKS, ECS, Lambda, RDS

WAF and Shield hardening, centralised secrets (Secrets Manager, KMS), security and performance alarms on a secure data platform.

DevSecOps

Security built into the CI/CD chain

SAST · quality · E2E

Static analysis, SonarQube and end-to-end tests built into the Jenkins pipelines, vulnerabilities fixed, artefacts managed, jobs parallelised.

A citizen services programme · Azure

End-to-end Azure environments

dev · test · sandbox

Three isolated environments from day one of the project, network security rules, real-time FinOps alerts, centralised secrets.

The platform’s delivery chain

As it exists in the repository, with the real status of each link.

  1. written, bench-tested

    Multi-architecture continuous integration

    On every push: secret scanning, test benches, dependency audit, page checks; then an image built for two processor architectures, and scanned.

  2. written, bench-tested

    Infrastructure as code

    DNS, TLS and firewall described as code, validated against the real provider, tested by a bench. Nothing is applied without a reviewed plan.

  3. written, bench-tested

    Kubernetes and GitOps

    The manifests live in the repository; a GitOps controller reconciles them read-only. Pinned versions, full rehearsal on a disposable machine.

  4. in service

    Alerts and probes

    An error workflow attached to every production workflow, three delivery attempts. Every morning, a probe compares declared state with real state.

  5. next step

    SLOs and error budget

    Numeric service objectives, read over thirty days, and burn-rate alerts on the error budget.

  6. in service

    Runbooks

    Executable procedures in the repository: install, switch over, restore, roll back, every step in order.

  7. in service

    Blameless post-mortems

    Every failure written up with the same template: problem, measurement, decision, result, what I would do differently. We look for a cause, never a person.

Readout

What the platform did over the last seven days

Three counters, read by a script from the database of the production workflow engine, at the time given below the figures. None is typed by hand.

11
active production workflows
read from the workflow engine
105
successful runs over seven days
success status, rolling window
9
failed runs over seven days
each one raised an alert

Last written: .

Going live

What the platform learned to do, in order

Every step ran for real before the next one.

  1. step 01

    One door, deterministic routing

    Everything enters through a messaging app. A rule, not a model, decides what is private, and private data never meets a model. The rest becomes a note or a task in a second brain.

  2. step 02

    The first autonomous production workflow

    Topic, writing, generation, checks, assembly, delivery: a complete chain with no human hand between the trigger and the deliverable.

  3. step 03

    One bot per role

    Capture keeps its door; conversation with an agent gets its own. Every use has exactly one door.

  4. step 04

    Production leaves the workstation

    Nothing depends on a computer being switched on anymore. Everything that produces lives on the server; the workstation is for trials and measurements.

  5. step 05

    Inference on an on-demand GPU

    The GPU load moves to a serverless service billed by the second. Measured cost per run: $0.041.

  6. step 06

    System state is generated

    A state page produced by script from declarations, a test bench that verifies it, a probe comparing declared and real.

  7. step 07

    Encrypted backup, secret guard, agent journal

    Secrets leave the machine encrypted every night; a guard blocks any key before it leaves the machine; every agent session is journaled.

  8. step 08

    Runs every day, seven days a week

    Delivery to every destination, metrics read back, an alert on every failure.

The platform

Explore it: every node is something that runs

Self-hosted infrastructure on a cloud server, on-demand GPUs, and nothing that exists in only one place. Select a node to open its details.

node

The production engine

daily workflows

Writing, generation, checks, assembly, delivery. Every run journaled, every failure alerted.

every run journaled · every failure alerted

generatecheckdeliver

A run in two phases

The real shape of a production workflow, redrawn with generic nodes: no vendor, no service, no identifier. Select a node to see what it does.

  • The shapePhase 1 prepares, drops and stops. Outside the platform, a scheduled agent generates; the connector checks every drop.
  • The wake-upOn the last drop, a webhook restarts the same workflow, once. No executor waits.
  • The resumeThe run state is read back: what is done is not redone.

Built for agents

A platform whose users are AI agents

The constraint is the same as for a team of developers: they must be able to act without me, and I must be able to verify what they did. It forced me to write explicit rules, generate the system state, cover everything with test benches, and journal every session. Here are the four mechanisms, with real excerpts.

01

The entry point

| when | what |
| by default, every session | AGENTS.md |
| first thing to open | system state, generated |
| on demand | naming rules, registry, decisions |

One tool-agnostic file every agent reads first, whatever the vendor. It tells nothing: it says where things are told, and what is not up for debate.

02

Generated state

project: the proof site
status: in production
triggered: on every build
→ state page regenerated, bench green

Written by a script from declarations, never by hand. A project that is not declared does not exist for the system.

03

Test benches

7 · The registry counts right
  ok  a number used twice is caught
  ok  a bare pipe in a cell stops the script
44 checks passed, 0 failed

The verdict is “0 failed”, never a total: totals change with tooling, and an expected number looks like a regression the day it moves.

04

Session journal

{ "duration_min": 45,
  "repos": { "governance": 1, "projects": 2 },
  "written_by": "journal_session",
  "entry": "agent-b" }

Every agent session leaves one line: its duration and the repositories it touched. Two agents from different vendors write to it in the same format.

What it taught me, and what it cost

01I measure before I fix

Two fixes went out before the signal had been measured. As soon as it was, the cause showed up in one pass. Since then, every fix starts with a measurement.

02Every service reports when it is not working properly

A service answered 200 for six weeks while skipping a step. It now reports that it is degraded, and that alert was checked in production.

03A failure count is written down with its time

44 of 49 in the evening, 46 of 51 the next day: the count was rising while I wrote the fix. I note the time of the reading and how fast it rises.

04The code checks what the model writes

After three fixes of the same kind in a prompt, I changed method: the model fills fields defined in advance, and the code checks each field before going on.

05The status page is produced by a script

Kept by hand, a status page becomes wrong within a few days. Produced by a script and checked by a test bench, it matches what is running every morning.

06Every workstream is listed in a registry

Three roadmaps sat in three folders, under three different names; finding the third meant searching the disk. There is now one registry, one line per workstream, updated the same day.

Case studies

Four problems, four measurements, four decisions

Same skeleton every time: the problem, the measurement, the decision, the result, and what I would do differently. No number appears without how it was measured.

2025 to 2026 · AWS

Cutting 65 to 75 % of a European aircraft manufacturer’s cloud costs, with zero incidents

65 to 75 % · 0 incidents

Problem

A secure data platform, an AWS bill growing faster than usage, and a production that could not be stopped.

Measurement
provisioning audit, service by service, over 30 days of CloudWatch metrics: real CPU, memory and IOPS against reserved
Decision

Reduce in two one-week steps, with alarm validation between them, rather than all at once. Rejected: a single-pass cut, and spot instances on ingestion workloads that could replay files.

Result
65 to 75 % reduction depending on the service, read on billing; 0 production incidents during and after the cut, read on the alarms
What I would do differently

Set threshold alarms before the cut rather than observe after, and automate over-provisioning detection instead of auditing by hand.

data

Private data never meets a model

0 private items sent to an API

Problem

An assistant that triages everything coming in, including voice messages and photos of invoices. Part of it must never leave the machine.

Measurement
every execution re-read from the database: the message, the class returned, the model called, the tokens; 24 real executions over one night
Decision

A deterministic rule, not a model, decides what is private. First a local model for private data; then, measured too slow to start, removed: private data is kept verbatim, with no model at all, and invisible to the assistant’s reading tools.

Result
0 private messages sent to an API, by construction; classification 8 of 8 correct on the measured batch
What I would do differently

Test the local model cold before installing it: 175 seconds of loading on the server is measured in ten minutes, not after two weeks.

FinOps

GPU inference leaves the workstation

$0.041 per run

Problem

A generative model running on a laptop’s GPU. Switched off or busy, the workflow waited 58 minutes for a result that never came, without raising an error.

Measurement
cost per run re-read from the service log, over 6 real jobs: $0.037 to $0.047
Decision

Inference on a serverless GPU billed by the second; post-processing, which is light computation, stays on the server. Rejected: a dedicated cloud GPU, paid for even when idle.

Result
$0.041 per run, about $6 a month for five runs a day; nothing depends on the workstation anymore
What I would do differently

Re-read the cost instead of copying it: a code comment announced $0.0675 for ten days after the real measurement.

reliability

The silent failure: 46 of 51 runs degraded, all returning HTTP 200

the failure is no longer silent

Problem

A post-processing service looked for a component at a hard-coded address. The component restarted, the address changed, and the service kept answering 200 while skipping the step.

Measurement
service log counted line by line: 46 of 51 jobs carried “unreachable”; three contradictory counts before matching the exact string, spaces included
Decision

The service states its condition: a health probe returns “degraded” with the reason, the report carries the post-processing status, the log writes it. The component is reached by name, no longer by address. Rejected: silently regenerating when the component is missing.

Result
deployed and observed: the probe returns degraded as long as the component is not published, and says so; the failure is no longer invisible
What I would do differently

Forbid any hard-coded address from the first version, and date every failure count: it was rising by two a day while I was writing.

Background

Four industries, four technology stacks

Four years with the same IT services employer, on assignment with industrial clients. Clients are named by industry.

  • A European aircraft manufacturer · secure data platform

    AWS costs reduced by 65 to 75 %; ingestion pipelines (Glue, EventBridge, S3); static analysis and quality gates in pipelines; WAF hardening, secrets and KMS; Terraform and CloudFormation; AI-augmented search on GCP.

  • A citizen services programme · Azure

    End-to-end Azure environments, network security rules, FinOps alerts.

  • An agricultural equipment manufacturer · dealer mobile apps

    Technical debt reduced by about 80 %, Cordova to Capacitor migration, mobile CI/CD, six quarterly store releases.

  • A commercial vehicle manufacturer · fleet monitoring

    AngularJS to Angular 16 migration in a France and India project of over one hundred people.

  • A food-service startup · apprenticeship

    Complete AWS infrastructure, serverless PWA, payments, iOS and Android releases.

Cloud and infrastructure as code

AWSAzureGCPTerraformOpenTofuCloudFormation

Kubernetes and containers

KubernetesEKSECSDockerHelmGitOps

CI/CD and supply-chain security

JenkinsGitLab CIAzure DevOpsArtifactorySASTSonarQubeCypressKMSWAF

Observability and FinOps

CloudWatchalarmsprobesAWS Budgetscost alerts

AI agents and automation

MCPagent SDKsn8nRAGdeterministic routingLLM FinOps

Languages

Python · FastAPIJavaScript · TypeScriptBashSQLC#Java

CV

The full CV

The complete version, on its own page, printable on two A4 pages.

Boubacar Soumaré

DevOps and Platform Engineer · AWS and Azure cloud, AWS certified

Toulouse, France · contact@boubacarsoumare.com · linkedin.com/in/boubacar-soumare

Profile

DevOps, Cloud and Platform engineer with four years of experience on critical industrial infrastructure (aerospace, agricultural machinery, commercial vehicles). Specialised in multi-service AWS and Azure environments, I industrialise CI/CD and DevSecOps chains (Jenkins, GitLab CI, Checkmarx) and drive cloud efficiency: a measured 65 to 75 % reduction in AWS costs at a European aircraft manufacturer, with no production interruption. With a software engineering background (Python and multi-language tooling), I treat infrastructure as modular, tested code (Terraform). Alongside, I run a self-hosted infrastructure platform in continuous 24/7 production (containers, AI automation, agents), with the same operational rigour.

Certifications

AWS Certified Cloud Practitioner (2024)

CKA, Certified Kubernetes Administrator (in preparation)

Education

MSc in application development, Epitech Toulouse (2022)

BSc in computer science, Université de Haute-Alsace, Mulhouse (2020)

Languages

French, native

Professional English (B2): four years in an international environment, daily meetings with teams in France and India

Professional experience

An IT services company · software engineer, then DevOps and cloud engineer
  • A European aircraft manufacturer · secure big data platform Aug 2025 → today
  • A citizen services programme · Azure environments May → Jul 2025
  • An agricultural equipment manufacturer · dealer network mobile apps Nov 2023 → Apr 2025
  • A commercial vehicle manufacturer · fleet monitoring 2022 → 2023
A food-service startup · full-stack developer, apprenticeship
  • A platform connecting restaurant owners and cooks
Automation and AI platform, in production
  • Designed, built and run alone, in 24/7 production.

Contact

Write to me

A project, a question about the platform, an idea to discuss: one email is enough, I reply.

contact@boubacarsoumare.comlinkedin.com/in/boubacar-soumare