|
KEY ACCOUNTABILITIES & ACTIVITIES
This section describes the principal outputs required from the job.
|
|
Key Accountabilities
|
Key Activities
|
- AI Model Deployment & Operations
|
- Deploy ML and LLM models into production environments using containers and/or cloud services in line with approved technical requirements.
- Configure deployment parameters, autoscaling, and runtime settings to support stable and scalable AI services.
- Monitor production deployments and support resolution of operational issues affecting model availability and performance.
|
- AI Performance & Resource Optimization
|
- Monitor and optimize AI model performance, latency, error rates, resource utilization, and associated infrastructure costs.
- Implement appropriate optimization techniques such as caching, pruning, quantization, resource quotas, and cost tuning.
- Configure dashboards and alerts to provide visibility into model health, resource consumption, and output quality.
|
- AI Solution Engineering & Integration
|
- Develop AI solutions based on models and capabilities produced by GenAI and research teams.
- Build service layers and APIs around AI models and package solutions for deployment using appropriate engineering frameworks.
- Integrate AI solutions with required data sources, identity services, logging capabilities, and enterprise systems in line with Elm standards.
|
- Model Lifecycle & Release Management
|
- Support model lifecycle management through model registries, versioning, controlled releases, and safe rollback mechanisms.
- Develop and maintain unit and integration tests to support reliable solution releases.
- Establish and maintain CI/CD pipelines to automate testing, deployment, and release activities for AI solutions.
|
- AI Evaluation & Validation
|
- Conduct automated and human-in-the-loop assessments to evaluate AI model accuracy, reliability, and output quality against agreed metrics.
- Monitor data and concept drift and support identification of model retraining or improvement requirements.
- Prepare evaluation results highlighting identified risks, limitations, performance gaps, and recommended actions.
|
- AI Infrastructure & Production Readiness
|
- Coordinate with infrastructure teams to provision required GPU/CPU, storage, networking, and cloud resources for AI workloads.
- Support the setup and integration of enabling technologies such as vector databases, feature stores, and model or prompt registries.
- Support backup, recovery, resilience, and production-readiness requirements for deployed AI solutions.
|
- Technical Delivery & ODC Collaboration
|
- Coordinate with ODC teams on development, testing, and technical delivery activities in accordance with agreed delivery and quality requirements.
- Review code, technical artifacts, and deliverables received from ODC teams to verify alignment with Elm’s technical and security standards.
- Support consistent engineering practices, templates, and technical documentation across delivery teams and centers.
|
- Technical Documentation & Knowledge Transfer
|
- Develop and maintain technical documentation, deployment guides, runbooks, configurations, and operational procedures for AI solutions.
- Document solution architecture, dependencies, configurations, and known operational considerations to support maintainability and continuity.
- Support knowledge transfer to relevant engineering, infrastructure, and operational teams to enable effective solution support.
|
- Policies, Processes & Procedures
|
- Follow all relevant departmental policies, processes, standard operating procedures, and instructions so that work is carried out in a controlled and consistent manner.
- Comply with all relevant safety, quality, and environmental management policies, procedures, and controls to ensure a healthy and safe work environment.
|
- Information Security
|
- Comply with all relevant information security practices and standards to ensure data integrity and confidentiality.
|