Req number:
R8164Employment type:
Full timeWorksite flexibility:
HybridWho we areCAI is a global services firm with over 9,000 associates worldwide and a yearly revenue of $1.3 billion+. We have over 40 years of excellence in uniting talent and technology to power the possible for our clients, colleagues, and communities. As a privately held company, we have the freedom and focus to do what is right—whatever it takes. Our tailor-made solutions create lasting results across the public and commercial sectors, and we are trailblazers in bringing neurodiversity to the enterprise.
Job Summary
The Senior AI Engineer is to design, build, integrate, and operate secure, scalable, and production-ready AI solutions within an enterprise environment.The role combines strong Python / .NET software engineering with generative AI engineering, AWS-native cloud architecture, enterprise system integration, observability, and application security. The Senior AI Engineer will provide technical leadership across the engineering lifecycle, from solution design and experimentation through implementation, production deployment, monitoring, and continuous improvement.
Job Description
We are looking for Senior AI Engineer to design, build, integrate, and operate secure, scalable, and production-ready AI solutions within an enterprise environment. The role combines strong Python / .NET software engineering with generative AI engineering, AWS-native cloud architecture, enterprise system integration, observability, and application security. The Senior AI Engineer will provide technical leadership across the engineering lifecycle, from solution design and experimentation through implementation, production deployment, monitoring, and continuous improvement. This position will be full-time and hybrid at Mandaluyong City.
What You'll Do
Translate business and product requirements into appropriate technical designs, implementation plans, and engineering tasks
Lead technical design reviews and contribute to architecture, security, data, and operational readiness assessments.
Evaluate technical options and provide recommendations based on feasibility, scalability, security, performance, cost, and maintainability
Identify technical dependencies, delivery risks, resource requirements, and architectural constraints early in the development lifecycle
Guide engineers in resolving complex technical issues involving AI models, application services, data pipelines, cloud infrastructure, and enterprise integrations
Support technical estimation, work planning, backlog refinement, and delivery prioritization
Promote the use of shared enterprise AI capabilities and reusable platform services rather than duplicating solution-specific implementations
Mentor junior and mid-level engineers through code reviews, design discussions, pair programming, and knowledge-sharing sessions.
Backend Software Engineering
Design and develop high-quality Python and/or .NET services, libraries, APIs, background workers, and data-processing components
Build modular, reusable, testable, and maintainable application components using established Python and/or .NET engineering practices
Develop synchronous and asynchronous services that support AI inference, document processing, data retrieval, workflow orchestration, and system integration
Implement appropriate exception handling, retry mechanisms, timeouts, circuit breakers, caching, rate limiting, and graceful degradation
Apply object-oriented, functional, domain-driven, and event-driven design approaches where appropriate
Develop automated unit, integration, contract, security, performance, and regression tests
Maintain clear technical documentation covering solution architecture, APIs, configuration, deployment, operations, and troubleshooting
Contribute to continuous integration and continuous delivery pipelines for automated testing, security scanning, deployment, and release management
Participate in code reviews and ensure that engineering work meets agreed quality, security, performance, and maintainability standards.
System Integrations
Design and implement secure integrations between AI solutions and enterprise applications, data platforms, document repositories, workflow systems, and external services
Develop and maintain REST, event-driven, messaging, streaming, batch, and file-based integration patterns
Build integrations using APIs, webhooks, message queues, event buses, managed file transfer, and other approved enterprise integration mechanisms
Implement authentication and authorization using enterprise identity standards such as OAuth 2.0, OpenID Connect, service identities, API credentials, and role-based access controls
Integrate AI solutions with structured and unstructured data sources while preserving source permissions, data classifications, and access-control requirements
Develop connectors for enterprise systems such as document management platforms, service management tools, data warehouses, databases, search platforms, and business applications
Define API contracts, data schemas, error-handling conventions, versioning strategies, and integration testing requirements
Coordinate with application owners and platform teams to resolve integration constraints, access requirements, service limits, and dependency timelines
Ensure that integrations are observable, resilient, idempotent where necessary, and designed to handle partial failures safely.
AWS-Native Cloud Engineering
Design and implement AI solutions using approved AWS-native cloud services and architectural patterns
Develop solutions using relevant services such as Azure OpenAI, Amazon Bedrock, AWS Lambda, Amazon ECS, Amazon EKS, Amazon API Gateway, Amazon S3, Amazon RDS, Amazon OpenSearch Service, Amazon EventBridge, Amazon SQS, Amazon SNS, AWS Step Functions, and AWS Secrets Manager
Implement cloud-native patterns for serverless processing, containerized workloads, event-driven architecture, workflow orchestration, batch processing, and API-based services
Design solutions that meet enterprise requirements for availability, scalability, resilience, performance, backup, disaster recovery, and cost management
Work with cloud infrastructure teams to define network connectivity, private endpoints, security groups, encryption, logging, and environment configurations
Contribute to infrastructure-as-code implementations using Terraform
Optimize cloud resource usage, model consumption, storage, data transfer, and compute costs
Support deployments across development, testing, staging, and production environments
Troubleshoot application, platform, network, permissions, capacity, and service-integration issues within AWS environments.
Generative AI Engineering
Design and build generative AI solutions using foundation models, large language models, embedding models, reranking models, and multimodal capabilities
Develop retrieval-augmented generation applications that combine enterprise content, search services, vector retrieval, metadata filtering, and generative models
Design prompt templates, system instructions, tool descriptions, response schemas, and conversation flows
Implement model routing, fallback, retry, timeout, caching, and rate-limiting mechanisms
Build AI agent and workflow capabilities that can select approved tools, retrieve information, invoke enterprise services, and complete controlled multi-step tasks
Develop document ingestion, parsing, chunking, metadata enrichment, embedding, indexing, retrieval, and citation-generation pipelines
Implement hybrid search approaches that may combine keyword search, semantic search, vector search, taxonomy, graph-based retrieval, and reranking
Work with Data Scientists and AI Engineers to evaluate models and AI responses using measurable criteria such as relevance, groundedness, accuracy, completeness, safety, latency, and cost
Develop automated and human-in-the-loop evaluation processes for prompts, models, retrieval strategies, and generated outputs
Implement safeguards against prompt injection, insecure tool invocation, sensitive-data exposure, hallucination, inappropriate content, and unauthorized information retrieval
Support AI red-teaming, adversarial testing, security testing, and responsible AI assessments
Monitor model behaviour, token consumption, response quality, retrieval performance, latency, failures, and operational cost
Document model limitations, solution assumptions, evaluation results, human oversight requirements, and appropriate-use conditions
Collaborate with product managers and business stakeholders to clarify requirements, intended outcomes, user expectations, and acceptance criteria
Work with solution and enterprise architects to ensure alignment with organizational architecture principles and technology standards
Partner with data engineers and data owners to address data quality, availability, lineage, classification, access, licensing, and refresh requirements
Coordinate with infrastructure, platform engineering, and site reliability teams to establish production environments and operational support models
Work closely with cybersecurity, data security, risk, legal, compliance, and responsible AI teams to implement required controls
Participate in agile ceremonies, technical workshops, architecture reviews, security reviews, and operational readiness assessments
Communicate technical risks, dependencies, design decisions, implementation trade-offs, and delivery progress clearly to technical and non-technical stakeholders
Support delivery partners and vendors by defining technical expectations, reviewing deliverables, and ensuring alignment with enterprise engineering standards
Contribute to engineering communities of practice, reusable technical assets, reference implementations, and internal knowledge repositories
Foster a collaborative engineering culture that encourages constructive challenge, early escalation, shared ownership, and continuous improvement
Implement application, infrastructure, model, retrieval, integration, and user-experience observability using Datadog
Configure structured logging, metrics, distributed tracing, dashboards, alerts, monitors, and service-level indicators
Instrument Python services to provide visibility into request processing, dependencies, database activity, external API calls, model invocations, and background jobs
Establish correlation across application logs, traces, infrastructure telemetry, model requests, and business transactions
Monitor AI-specific operational metrics such as model latency, token usage, inference failures, retrieval quality, empty results, fallback rates, and cost
Define meaningful alerts that support early detection while minimizing unnecessary operational noise
Support incident investigation, root-cause analysis, performance tuning, and post-incident improvement activities
Ensure that logs and telemetry are designed to avoid exposing credentials, confidential data, personal information, prompts, or model responses without appropriate controls
Develop operational dashboards for engineering teams, service owners, product managers, and support functions
Contribute to service-level objectives, availability targets, performance baselines, and capacity planning
Embed secure software development practices throughout the design, development, testing, deployment, and operation of AI solutions
Use Wiz to identify and address cloud security risks, exposed resources, configuration weaknesses, excessive permissions, vulnerable workloads, and security posture issues
Use Snyk to identify and remediate vulnerabilities in open-source dependencies, application code, container images, and infrastructure-as-code
Integrate security scanning and policy checks into continuous integration and continuous delivery pipelines
Review and remediate software composition analysis, static application security testing, container scanning, infrastructure-as-code scanning, and cloud posture findings
Apply secure coding practices to prevent common vulnerabilities involving injection, insecure deserialization, authentication, authorization, sensitive-data exposure, and server-side request forgery
Securely manage application secrets, API keys, certificates, database credentials, and model-provider credentials
Implement appropriate encryption, access controls, audit logging, input validation, output filtering, and data-handling controls
Assess third-party Python packages, AI frameworks, models, containers, and external services before adoption
Work with cybersecurity teams to prioritize findings based on exploitability, business impact, data sensitivity, and production exposure
Track security findings, exceptions, compensating controls, and remediation actions through till closure
Support threat modelling and security assessments for AI applications, model integrations, retrieval pipelines, and agentic workflows.
Ensure that AI services are production-ready, supportable, observable, secure, and resilient before release
Contribute to deployment plans, rollback procedures, operational runbooks, support documentation, and post-deployment validation
Support the preparation of technical evidence required for architecture, cybersecurity, operational risk, and Change Approval Board reviews
Participate in production deployments and provide technical support during agreed implementation and stabilization periods
Investigate production incidents and implement corrective and preventive improvements
Manage technical debt and ensure that deferred engineering or security actions are recorded, prioritized, and remediated
Support capacity planning, performance testing, resilience testing, backup validation, and disaster recovery exercises
Continuously improve reliability, performance, security, developer productivity, and operational efficiency
Reliable and maintainable Python AI services delivered in accordance with agreed engineering standards
Secure, scalable, and observable AI solutions operating successfully in AWS environments
Effective integration with enterprise applications, data sources, identity services, and shared platforms
Early identification and resolution of technical, security, data, and operational risks
Measurable improvements in AI response quality, retrieval performance, service reliability, and operational cost
Timely remediation of security findings identified through Wiz, Snyk, and related engineering security controls
Reduced production incidents through effective testing, observability, resilience engineering, and operational readiness
Increased reuse of shared AI components, libraries, integration patterns, and platform capabilities
Strong engineering collaboration, knowledge transfer, and development of less-experienced team members
Successful transition of AI capabilities from experimentation into governed, secure, and supportable enterprise services
What You'll Need
Required:
Significant professional experience developing production-grade applications using Python
Demonstrated experience designing and delivering enterprise AI, machine learning, data, or cloud-native solutions
Strong understanding of Python and/or .NET frameworks and libraries used for APIs, asynchronous processing, data engineering, testing, and AI development
Practical experience developing generative AI, retrieval-augmented generation, intelligent search, or AI agent solutions
Hands-on experience with AWS-native cloud services and cloud-native architectural patterns
Experience designing and integrating REST APIs, event-driven services, databases, enterprise applications, and data platforms
Strong understanding of software architecture, distributed systems, microservices, application security, testing, and continuous delivery
Experience implementing observability using Datadog or a comparable enterprise observability platform
Experience using cloud and application security tools such as Wiz and Snyk
Experience with container technologies and orchestration platforms such as Docker, Amazon ECS, or Amazon EKS
Experience with automated testing, source control, code review, dependency management, and CI/CD pipelines
Strong analytical, troubleshooting, technical writing, and stakeholder communication skills
Demonstrated ability to lead technical discussions, mentor engineers, and influence engineering decisions without relying solely on formal authority
Preferred:
Experience delivering AI solutions in a regulated, financial, public-sector, or risk-sensitive enterprise environment
Experience with Amazon Bedrock, vector databases, search platforms, model gateways, and AI evaluation frameworks
Familiarity with responsible AI principles, model risk, bias assessment, explainability, human oversight, and AI governance
Experience implementing model routing, AI guardrails, prompt security, retrieval security, and agent tool controls
Familiarity with infrastructure as code, DevSecOps, MLOps, LLMOps, platform engineering, and site reliability engineering
Experience with enterprise identity platforms, private cloud connectivity, API management, and role-based access control
Relevant certifications or formal education in software engineering, cloud architecture, cybersecurity, data engineering, artificial intelligence, or a related discipline
Advanced Python and/or .NET engineering
Generative AI and retrieval engineering
AWS cloud architecture
Enterprise system integration
Secure software development
Technical leadership and mentoring
Observability and operational troubleshooting
Architecture and design thinking
Structured problem-solving
Technical risk management
Stakeholder communication
Team collaboration
Continuous learning and improvement
Strong problem-solving abilities and attention to detail
Strong communication skills to articulate technical dependencies, delivery risks, resource requirements, and architectural constraints to product managers and technical delivery manager
Reporting and Coordination
Physical Demands
Ability to safely and successfully perform the essential job functions
Sedentary work that involves sitting or remaining stationary most of the time with occasional need to move around the office to attend meetings, etc.
Ability to conduct repetitive tasks on a computer, utilizing a mouse, keyboard, and monitor
Reasonable accommodation statement
If you require a reasonable accommodation in completing this application, interviewing, completing any pre-employment testing, or otherwise participating in the employment selection process, please direct your inquiries to application.accommodations@cai.io or (888) 824 – 8111.
