Key Responsibilities
Generative AI and LLM Application Development
Design and develop enterprise Generative AI applications using commercial and open-source Large Language Models.
Build AI assistants, enterprise search applications, document intelligence solutions, conversational systems and workflow automation tools.
Integrate LLMs with existing applications, APIs, databases, content repositories and enterprise systems.
Evaluate and select suitable foundation models based on accuracy, latency, cost, privacy and business requirements.
Develop reusable AI services, APIs and components for multiple business applications.
Retrieval-Augmented Generation
Design and implement end-to-end Retrieval-Augmented Generation pipelines.
Connect LLM applications with private company data, documents, databases and knowledge repositories.
Build document ingestion, cleaning, chunking, embedding, indexing and retrieval workflows.
Implement semantic search, hybrid search, metadata filtering, query rewriting, reranking and context management.
Improve retrieval accuracy, answer groundedness and source attribution.
Develop evaluation frameworks to measure retrieval relevance and response quality.
Implement access controls to ensure that users can retrieve only authorized enterprise information.
Vector Database Design and Management
Design, implement and optimize vector-search solutions using Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS, Elasticsearch, OpenSearch, Azure AI Search or PostgreSQL with pgvector.
Select suitable embedding models, indexing strategies and similarity metrics.
Optimize vector databases for retrieval speed, accuracy, scalability and cost.
Manage data ingestion, index updates, metadata, document deletion and versioning.
Monitor vector-search performance and troubleshoot retrieval-quality issues.
Implement hybrid retrieval combining semantic and keyword-based search.
Model Fine-Tuning and Optimization
Fine-tune open-source models such as Llama, Mistral, Gemma, Qwen or similar models using domain-specific datasets.
Apply parameter-efficient fine-tuning techniques such as LoRA, QLoRA and PEFT.
Prepare, clean and validate training, evaluation and instruction datasets.
Conduct model benchmarking, hyperparameter tuning and experiment tracking.
Compare prompt engineering, RAG and fine-tuning approaches to select the most appropriate solution.
Optimise models through quantisation, batching, caching and other inference techniques.
Evaluate models for accuracy, hallucination, bias, safety and domain suitability.
Deployment, MLOps and Scaling
Deploy AI and machine learning applications across cloud, containerised and on-premises environments.
Build scalable model-serving APIs using Python, FastAPI, Flask or similar frameworks.
Use Docker, Kubernetes and cloud services to deploy and scale AI workloads.
Implement automated CI/CD pipelines for AI applications and model releases.
Design systems capable of supporting growth from small user groups to thousands of concurrent users.
Implement caching, load balancing, asynchronous processing, rate limiting and autoscaling.
Monitor latency, throughput, token usage, infrastructure utilization, API errors and inference costs.
Optimise cloud and compute costs without compromising application performance or reliability.
AI Evaluation and Observability
Build automated and human-led evaluation frameworks for LLM applications.
Measure answer accuracy, relevance, groundedness, retrieval quality, latency and cost.
Create test datasets covering common queries, edge cases and failure scenarios.
Implement tracing and observability for prompts, retrieval steps, model calls and final responses.
Identify hallucinations, retrieval failures and changes in model behavior.
Conduct regression testing following model, prompt, data or architecture changes.
Security and Responsible AI
Protect confidential and personally identifiable information used by AI applications.
Implement authentication, authorization, encryption and role-based access controls.
Address prompt injection, data leakage, insecure output handling and other LLM security risks.
Apply responsible AI practices relating to fairness, transparency, safety and explainability.
Maintain documentation and audit trails for models, datasets, prompts and deployments.
Technical Leadership
Convert business requirements into scalable AI solution designs.
Participate in architecture discussions, code reviews and technical decision-making.
Guide junior engineers and share best practices for GenAI application development.
Work closely with client stakeholders, product managers, data engineers and software-development teams.
Prepare technical documentation, architecture diagrams, deployment guides and operating procedures.
Required Skills and Experience
3 to 8 years of experience in machine learning, artificial intelligence, data science, software engineering or a related field.
Strong hands-on experience building Generative AI or LLM-based applications.
Practical experience designing and implementing RAG pipelines.
Experience with vector databases or vector-search platforms.
Strong proficiency in Python.
Experience with frameworks such as LangChain, LlamaIndex, Haystack, Semantic Kernel or similar platforms.
Experience with LLM APIs such as OpenAI, Azure OpenAI, Amazon Bedrock, Google Gemini or Anthropic.
Experience working with open-source models and the Hugging Face ecosystem.
Understanding of embeddings, tokenization, transformers, attention mechanisms and context windows.
Experience developing and integrating REST APIs.
Strong knowledge of Git, software-development practices and application testing.
Good understanding of cloud architecture, deployment and production monitoring.
Strong analytical, problem-solving and communication skills.
Ability to work independently in a remote environment aligned with UK business hours.
Preferred Skills
Experience fine-tuning open-source Large Language Models.
Knowledge of LoRA, QLoRA, PEFT, quantization and distributed training.
Experience with PyTorch, TensorFlow or JAX.
Familiarity with MLflow, Weights & Biases, LangSmith or similar experiment-tracking and observability tools.
Experience with vLLM, Hugging Face Text Generation Inference, NVIDIA Triton or similar model-serving frameworks.
Knowledge of Docker, Kubernetes and infrastructure-as-code tools.
Experience with Azure, AWS or Google Cloud AI and machine learning services.
Understanding of GPU infrastructure, memory management and model inference optimization.
Experience working with structured, unstructured and multimodal enterprise data.
Knowledge of SQL, NoSQL databases and data-engineering pipelines.
Experience in regulated domains such as healthcare, legal services, financial services or engineering will be advantageous.
Educational Qualifications
Bachelor's or Master's degree in Computer Science, Artificial Intelligence, Machine Learning, Data Science, Engineering or a related discipline.
Candidates with equivalent practical experience and strong production-level AI engineering expertise may also be considered.
What Makes This Opportunity Attractive?
Permanent work-from-home arrangement.
Opportunity to work directly with a UK-based client.
Hands-on involvement in enterprise Generative AI and LLM projects.
Work across RAG, vector search, model fine-tuning, deployment and MLOps.
Opportunity to build AI solutions that move from prototype to production.
Both regular full-time and freelance engagement options.
Exposure to international stakeholders and complex business applications.
About itForte
itForte is a leading specialist IT recruitment company connecting highly skilled technology professionals with clients across the US, UK, Europe and Japan.
We specialize in talent acquisition across Artificial Intelligence, Cybersecurity, Cloud Technologies, Enterprise Applications, and Product and Software Engineering.
Professionals with strong Generative AI, machine learning and production engineering experience are invited to apply.