Job purpose
We are seeking a highly skilled Cloud AI Engineer to design, build, and operate the AI systems and agentic applications that power our products and client engagements. You will leverage your expertise in Google Cloud Platform, Python, and modern LLM frameworks to take generative AI solutions from prototype to production — building retrieval systems, autonomous agents, and the evaluation and monitoring infrastructure that keeps them reliable. Your work will directly shape how our organisation and our clients put AI into the hands of real users.
Duties and responsibilities
Agentic AI & LLM Application Development
- Design, build, and deploy production-grade AI agents and multi-agent workflows using frameworks such as Google's Agent Development Kit (ADK), LangChain, or equivalent.
- Implement tool use, function calling, and structured outputs, integrating agents with internal APIs, databases, and third-party services (including via the Model Context Protocol).
- Develop and iterate on prompts, context strategies, and orchestration logic; establish version control and review practices for prompts as first-class artefacts.
- Apply appropriate guardrails, input/output validation, and human-in-the-loop checkpoints for agents that take consequential actions.
- Evaluate build-versus-buy trade-offs across model providers, frameworks, and managed services, and make pragmatic recommendations.
Retrieval-Augmented Generation & Knowledge Systems
- Build and maintain RAG pipelines end to end: ingestion, parsing, chunking, embedding, indexing, retrieval, re-ranking, and grounded generation.
- Select and operate vector stores appropriate to the workload (e.g. Vertex AI Vector Search, pgvector on Cloud SQL or AlloyDB, or equivalent).
- Implement hybrid and metadata-filtered retrieval strategies, and tune them against measured retrieval quality rather than intuition.
- Design document processing workflows across heterogeneous and messy source formats, including incremental refresh and deletion handling.
Cloud Engineering & Deployment
- Build, containerise, and deploy AI services and APIs on Google Cloud Platform, primarily using Cloud Run, Cloud Functions, and Vertex AI.
- Design and manage supporting cloud infrastructure — Cloud SQL, BigQuery, Cloud Storage, Pub/Sub, Artifact Registry, Secret Manager — with sound security choices around networking, IAM, and service accounts.
- Manage infrastructure as code (e.g. Terraform) and keep environments reproducible.
- Optimise for cost, latency, and throughput, including caching, batching, streaming responses, and right-sizing model selection to the task.
Data Engineering & Integration
- Build and maintain the data pipelines that feed AI systems, spanning batch and streaming ingestion from APIs, databases, and file sources.
- Write and optimize SQL against BigQuery and relational databases for both application queries and analytical workloads.
- Model and manage the operational data layer (Cloud SQL, Firestore, or similar) supporting agent state, session history, and application data.
- Implement data quality checks, schema management, and lineage where it materially affects downstream AI behaviour.
Evaluation, Monitoring & Observability
- Design and implement evaluation frameworks for LLM and agent systems, including golden datasets, offline eval suites, LLM-as-judge scoring, and regression testing across prompt and model changes.
- Instrument applications with tracing and structured logging (e.g. OpenTelemetry, Cloud Trace, Langfuse, Arize Phoenix, or equivalent) to make agent behaviour debuggable.
- Build monitoring and alerting for quality, latency, error rates, token consumption, and cost, and act on degradation proactively.
- Track groundedness, hallucination, and refusal behaviour in production, and close the loop between observed failures and system improvements.
- Investigate and troubleshoot AI-related incidents in production, including non-deterministic and hard-to-reproduce failures.
Python & Software Engineering Practice
- Write clean, tested, maintainable Python; contribute to shared libraries and internal tooling.
- Work within Git-based collaborative development practices, including branching strategies, pull requests, and code review.
- Apply sound API design, dependency management, and packaging practices to AI services.
- Balance rapid prototyping with the discipline required to make prototypes production-ready.
Documentation & Enablement
- Maintain clear documentation for AI systems, architectures, evaluation results, and operational runbooks.
- Communicate capabilities, limitations, and risks of AI systems honestly to technical and non-technical stakeholders.
- Support colleagues and clients in adopting the systems you build.
Qualifications
Essential:
- Proven experience building and deploying production applications on Google Cloud Platform, particularly Cloud Run, Cloud SQL, and Vertex AI.
- Strong Python proficiency, with demonstrable experience building services and applications.
- Hands-on experience designing and shipping LLM-powered systems, including at least one production RAG or agentic application.
- Practical familiarity with modern agent frameworks and orchestration patterns (ADK, LangGraph, LlamaIndex, or comparable).
- Solid SQL skills, including experience with BigQuery at scale.
- Experience with containerization (Docker) and deploying containerised workloads.
- Proficiency with version control systems and collaborative Git-based development practices.
- Demonstrated experience evaluating and monitoring AI systems in production, beyond manual spot-checking.
Preferred:
- Direct experience with Google's Agent Development Kit (ADK) and Vertex AI Agent Engine.
- Google Cloud Platform certification (e.g. Professional Machine Learning Engineer, Professional Cloud Developer, or Professional Data Engineer).
- Experience with infrastructure as code (Terraform) and CI/CD tooling.
- Experience with the Model Context Protocol (MCP) and building or consuming MCP servers.
- Workflow orchestration experience (Airflow / Cloud Composer, Dataflow, Kubeflow, or similar).
- Experience with LLM observability and evaluation tooling (Langfuse, Arize Phoenix, Weights & Biases, Vertex AI Evaluation, or equivalent).
- Frontend or full-stack experience sufficient to build usable interfaces over AI services.
- Experience with traditional ML and statistical modelling, and comfort collaborating with data scientists on hybrid systems.
- Familiarity with data visualisation tools (e.g. Looker, Data Studio, PowerBI).
- Client-facing or consulting experience.
Additional Skills:
- Strong problem-solving and analytical skills, with sound judgement about which problems warrant AI and which do not.
- Excellent communication and collaboration skills, including with non-technical stakeholders.
- Comfort operating amid rapid change in tooling and model capabilities, with the discipline to distinguish durable patterns from hype.
- Ability to learn new technologies quickly and adapt to changing requirements.
- Security- and privacy-conscious approach to handling data in AI systems.
Working conditions
- Competitive salary and benefits package.
- Opportunities for professional development and growth.
- Work in a dynamic and innovative environment.
- Remote work from home using personal equipment.
Physical requirements
None