Octus
Octus is a leading global provider of credit intelligence, data, and analytics. Since 2013, tens of thousands of professionals across hedge fund, investment banking, management consulting, and law firm verticals have come to rely on Octus to make better, faster, and more confident decisions in pace with the fast-moving credit markets. For more information, visit: https://octus.com/
Working at Octus
Octus hires growth-minded innovators and trailblazers across the globe to drive our business and culture. Our core values – Action Oriented, Customer First Mindset, Effective Team Players, and Driven to Excel – define an organizational ethos that’s as high-performing as it is human. Among other perks, Octus employees enjoy competitive health benefits, matched 401k and pension plans, PTO, generous parental leave, gym subsidies, educational reimbursements for career development, recognition programs, and much more. Role
As an AI Engineer focused on CreditAI, our flagship GenAI product, you will own complex technical problems across the full AI stack — designing distributed systems, orchestrating multi-agent workflows, and ensuring production reliability at scale.
Responsibilities
- Design and implement multi-agent and agentic orchestration frameworks using agent SDKs such as the Claude Agent SDK, Google ADK, or AWS AgentCore, incorporating tools, external data sources, memory, and state management
- Build and maintain MCP servers and integrations to extend AI system capabilities with structured tool use and external context
- Build and optimize RAG pipelines including embedding strategies, vector database, retrieval quality tuning, and cost-aware ingestion design
- Integrate with managed LLM services across cloud providers to support diverse deployment and cost optimization strategies.
- Fine-tune, optimize, and deploy open-source deep learning models for production use cases, leveraging GPU infrastructure for training and inference
- Apply systems thinking to design and optimize AI and LLM systems, balancing quality, scalability, latency, cost, and operational complexity, while implementing efficiency improvements using model selection, prompt design, batching, caching, and retrieval strategies.
- Design and implement automated evaluation frameworks to assess LLM system quality, accuracy, and performance across production workloads
- Apply reinforcement learning techniques (e.g., RLHF, RLAIF) to improve model alignment and task-specific performance
- Architect and manage high-throughput, real-time data pipelines using Kafka
- Design, deploy, and scale production AI services on AWS (Batch, Lambda, ECS, S3, etc), applying modern containerization, CI/CD, and infrastructure-as-code practices