← Sottava, jobs the hour they open
1 mo agofound 2 h ago
[VMT] Platform AI Software Engineer
What the posting is about
Own the trust layer of an AI copilot for process engineers, ensuring agent reliability and user trust. Investigate misbehavior, build evaluation harnesses, add reliability mechanisms, and diagnose performance issues in a large Python codebase. Collaborate with a global team using feature flags and async reviews.
Read out of the posting
Levelmid
Experience asked3+ years
EmploymentFull time
LocationBuenos Aires, Argentina
RemoteNot stated
Visa sponsorshipNot stated
SalaryNot published, and most postings do not
Posted2026-09-04
Found viasmartrecruiters, direct from their system
We saw it 1 month after it went up.
The posting, as the company wrote it
Employment: Full-time
Experience level: Mid-Senior Level
Job Description
Project - the aim you'll have
Our client builds an AI copilot for process engineers in oil refineries and chemical plants: a natural-language interface where engineers ask questions about live plant data — equipment, sensor tags, process trends — and get grounded, chart-backed answers. The users are experienced engineers who are rightly skeptical of AI: in this domain, a fabricated number or a silent wrong assumption has real cost. The product wins or loses on whether the agent can be trusted.
This role owns the trust layer of that agent inside a large, active Python codebase. It is not feature work with an LLM endpoint bolted on. The work is the mechanics of agent reliability: making the agent say "I don't know" instead of inventing, surfacing every assumption it makes so the user can correct it, grounding every claim in actual data, holding output quality through model migrations, and keeping latency acceptable while doing all of the above.
To make the day-to-day concrete, this is what the engineer currently in this seat shipped in the last four months (all of it flag-gated, in small PRs, reviewed async daily by a team spread across the US and Australia):
- An assumption auditor: detects the silent assumptions the agent makes when answering (which equipment, which time window), validates them via multi-draw consensus, and surfaces them in the UI as correctable chips — the engineer can fix an assumption and rerun the analysis.
- A grounding auditor that catches reports fabricated from empty data feeds before they reach the user.
- An adversarial reviewer sidecar that critiques generated charts for correctness before display.
- Successive frontier-model evaluations (loop behavior, directive adherence, regression on a replay harness) that decided when to flip the product's default model — including, twice, deciding NOT to flip.
- Hardening of a plant-exploration tool against hallucinating structure that the data does not support.
- A latency fix: a narration side-loop was inflating query response times; capped it and made it best-effort.
- A concurrency fix making a shared data-reset path atomic, eliminating intermittent production read errors.
If reading that list is more interesting to you than building another CRUD feature, this role is for you.
Qualifications
Expectations - the experience you need
Strong Python in large, shared, evolving backend codebases: you will work daily in code you didn't write, alongside people committing to it every day.
You have shipped an LLM-based feature to production AND built an evaluation that changed a real decision (a model choice, a prompt rollback, a killed feature).
Production debugging from symptom to confirmed root cause: latency spikes, concurrency errors, failures that produce no log line.
Prompt work treated as engineering: measured adherence, regression testing against a fixed case set — not vibes.
Comfort with feature-flag discipline and staged rollouts (default-off, soak, flip), small PRs, and mostly-async collaboration across US and Australia time zones.
High autonomy: problems arrive ambiguous ("the agent feels slow", "the engineers don't trust the numbers") and you turn them into scoped, verifiable fixes without waiting for a spec.
Direct, precise written English.
Nice to have
GCP (Vertex AI in particular); AWS/Azure acceptable.
Observability tooling (tracing, structured logging, latency percentiles you actually watched).
Experience with charting/plotting pipelines (matplotlib or similar) feeding a UI.
Industrial, process, or time-series data domain experience.
Heavy AI-tooling development workflow (Claude Code or similar) — the team works this way.
What you will do 
Investigate agent misbehavior reported from real customer plants and turn each case into a diagnosis, a fix, and a regression test.
Build and extend the evaluation harnesses that gate prompt changes and model migrations.
Add reliability mechanisms to the agent: assumption surfacing, grounding checks, output-quality reviewers.
Diagnose and resolve cross-cutting performance and concurrency issues.
Raise code quality in the areas you touch, within the team's review conventions.
Our Benefits  
Educational resources
Flexible schedule and Work From Anywhere
Referral Program
Supportive and chill atmosphere
We are accepting applications from LATAM countries
 
Company Description
We are Software Mind, an awesome team of engineers who are ready to ramp up any top-notch company’s projects! Our aim? To always be one step ahead. Become part of a multicultural company in constant growth with an excellent work environment certified by Great Place To Work!
Also open at Software Mind
Why this page exists
We read companies’ own hiring systems every hour, 1,182 of them, and show a job the hour it opens instead of when a job board gets around to indexing it. We saw it 1 month after it went up.
The feed is free. No card, no trial to expire.