CF CloudFrame Job Scanner Open Dashboard →
Verified Active Opening

AI Engineer (Managed Services)

Avepoint • Singapore

Job Description

<p>We are looking for a highly skilled AI Engineer specializing in Large Language Models (LLMs) and Agentic AI. You will architect, build, and deploy production-grade LLM applications — from intelligent knowledge bases and RAG systems to autonomous multi-agent workflows. You will work hands-on with open-source Chinese and international LLMs (DeepSeek, Qwen, Kimi, etc), implementing everything from model deployment and inference optimization to prompt engineering and agent orchestration. This is a builder role for someone who thrives at the intersection of research and engineering.</p> <p><strong>KEY RESPONSIBILITIES</strong></p> <p><strong><em>LLM Application Development</em></strong></p> <ul> <li>Design and develop enterprise LLM-powered applications: intelligent Q&A systems, enterprise knowledge base assistants, AI copilots, document analysis tools, and automated customer service agents.</li> <li>Architect and implement end-to-end RAG (Retrieval-Augmented Generation) systems: document parsing and chunking (recursive, semantic, agentic), embedding generation (BGE, M3E, GTE), vector retrieval (dense + sparse hybrid search), reranking (bge-reranker, Cohere Rerank), and response synthesis with source attribution.</li> <li>Develop and optimize Prompt Engineering strategies: chain-of-thought, tree-of-thought, few-shot prompting, structured output parsing (JSON mode / Pydantic), prompt templates (LangChain/LangSmith), and prompt version management.</li> <li>Knowledge in harness engineering, context management in ensuring LLM interactions and or AI agents reliable and deterministic.</li> </ul> <p> </p> <p> </p> <p><strong><em>AI Agent & Multi-Agent Systems</em></strong></p> <ul> <li>Design and build AI Agent systems using ReAct, Plan-and-Execute, Reflection, and multi-agent collaboration patterns.</li> <li>Implement Function Calling and tool-use capabilities, enabling agents to interact with external APIs, databases, and enterprise systems.</li> <li>Develop multi-agent orchestration using LangGraph, AutoGen, CrewAI, and other agent frameworks to solve complex enterprise tasks through agent collaboration.</li> <li>Design MCP (Model Context Protocol) integrations for standardized LLM tool interoperability.</li> </ul> <p><strong><em>Open-Source LLM Deployment & Optimization</em></strong></p> <ul> <li>Deploy and optimize latest version of open-source Chinese LLMs: DeepSeek, Qwen, and Kimi for on-premise and private cloud environments.</li> <li>Implement model inference optimization: quantization (GGUF/llama.cpp, GPTQ, AWQ, AutoAWQ, FP8/INT8), KV Cache optimization, continuous batching (vLLM, TensorRT-LLM, TGI, SGLang), speculative decoding, and tensor parallelism for high-throughput serving.</li> <li>Build and maintain model serving infrastructure using vLLM, TensorRT-LLM, Text Generation Inference (TGI), Ollama, Xinference, and SGLang; configure GPU resource scheduling with Kubernetes + GPU operators. AI gateway tools for routing, model tracking and load balancing such as TrueFoundry, Kubeflow, LiteLLM or Ray for heavy deep learning.</li> </ul> <p><strong><em>Model Fine-Tuning & Customization</em></strong></p> <ul> <li>Implement efficient fine-tuning pipelines using LoRA, QLoRA, DoRA, and full-parameter fine-tuning on proprietary domain-specific datasets.</li> <li>Prepare and curate instruction-following datasets, RLHF/RLAIF datasets, and evaluation benchmarks for domain adaptation.</li> <li>Evaluate fine-tuned models using automated benchmarks and LLM-as-a-Judge methodologies.</li> </ul> <p>

Job Reference ID: CF-158125 • Posted on CloudFrame Job Scanner