Transcript
When we first introduced InstrucLab, our goal was clear—making AI customization approachable and easy to try. And it worked. Developers got hands-on quickly, experimenting with fine-tuning and seeing the potential of bringing their data into large-language models. But as adoption grew, enterprises showed us the real challenge. AI training isn't always simple. Every use case has quirks. Business data looks different across industries. And making models truly useful requires expertise. Data scientists who understand how to clean the data, generating training examples, and choose the right training method. Push-button simplicity alone won't get you to production. That's why Red Hat AI model customization experience evolved beyond InstrucLab. Instead of one monolithic workflow, we now provide a modular architecture—a supported build of Python packages designed for every stage of your journey—DocLink for data processing, SCGHub for synthetic data generation, and a training hub for fine-tuning and continual learning. Because these components are modular, you can use them independently or connect them end-to-end. And with the supported cookbooks and notebooks, it's easy to start small, then scale the same workflows into production on OpenShift AI. Let's get into it. Let's first talk about DocLink. It's the number one open-source repository for document intelligence. DocLink lets you pre-process and structure enterprise documents with confidence—whether it's PDFs, HTML, Markdown, or Office files. And it's not just about local experimentation. With Red Hat's supported build, DocLink integrates directly into Kubeflow pipelines so you can process documents at scale, powering applications like Enterprise Search, RAC pipelines, and compliance workflows. Next, let's talk about Synthetic Data Generation Hub, a framework for building synthetic data pipelines. You can mix and match LLM-powered and traditional blocks, compose and orchestrate flows from simple transforms to multistage pipelines, easily extend and customize, create your own blocks, and plug into existing flows with minimal code. With SCGHub, synthetic data pipelines are modular, transparent, and production-ready. And to bring it all together, Training Hub delivers a stable, consistent interface for common training algorithms. Your teams gain access to the latest training methods, while Red Hat ensures API stability and enterprise support. Training Hub supports supervised fine-tuning from InstructLab training, orthogonal subspace learning for large language models, a new continual post-training algorithm, full compatibility with InstructLab's original multiphase pipeline, and integration with the latest open-source models like GPT-OSS for both SFT and post-training continual learning. To make this journey even easier, Red Hat AI provides supported cookbooks and examples that guide the customers through the full workflow, using DocLink for enterprise document intelligence, SCGHub for generating high-quality synthetic data training datasets, and Training Hub for fine-tuning models on that data. And that is why fine-tuning matters. Large language models today do not see the data that truly matters to an enterprise. Your internal documents, business processes, and domain expertise are invisible to general-purpose models out of the box. Fine-tuning is what bridges the gap, making models contextually relevant, accurate, and valuable for your teams and customers. We're not oversimplifying the problem. Instead, we give you enterprise-ready building blocks, flexible enough for experimentation and reliable enough for production. Your data scientists and engineers bring the expertise, and Red Hat provides the platform to make models smarter, faster, and more scalable. With Red Hat AI, fine-tuning is simplified, not by hiding complexity, but by embracing it. From preprocessing with DocLink to synthetic data pipelines with SCGHub to force-training algorithms with Training Hub and scalable deployment on OpenShift AI, now you have everything you need to connect your data to models.