Skip to main content
eScholarship
Open Access Publications from the University of California

UC Irvine

UC Irvine Electronic Theses and Dissertations bannerUC Irvine

Efficient Post-Training for Large Language Models and Agentic Systems

Creative Commons 'BY-NC-SA' version 4.0 license
Abstract

Large Language Models (LLMs) have shifted modern AI toward pretrained foundation models, where broad capabilities are acquired during pretraining and then specialized through post-training. While frontier-scale pretraining remains concentrated among a small number of organizations, the practical frontier for most researchers and product teams is post-training: adapting an existing backbone to new domains, user needs, deployment constraints, and evolving requirements. However, post-training is often costly and iterative in practice. Achieving high-quality adaptation typically requires substantial compute, careful data curation, repeated hyperparameter exploration, and rigorous evaluation to avoid regressions. These costs compound because post-training is rarely performed once: it is routinely repeated across domains and users, revisited after product changes, and re-run as deployment conditions shift. The resulting gap between the availability of strong pretrained backbones and the feasibility of repeated customization motivates the central theme of this thesis: efficient post-training.This thesis investigates efficient post-training for large language models and agentic systems as a lifecycle property, with the goal of reducing adaptation cost while preserving model capability, robustness, and alignment. We operationalize efficiency across the full post-training loop by considering not only the cost of gradient updates, but also the end-to-end expenses that dominate real deployments: trainable-state size, wall-clock training time, memory footprint, communication bandwidth, context growth, rollout volume, and evaluation overhead. A key premise is that no single technique universally optimizes all efficiency dimensions; instead, practical progress comes from co-design across algorithmic choices and system constraints. Accordingly, this thesis develops methods that (i) reduce the number of trainable parameters and stored specialization state, (ii) reduce communication and computation overhead per iteration in distributed or privacy-constrained settings, and (iii) improve the impact of each token and each interaction trajectory during training, especially for long-horizon agentic tasks.To structure the landscape, we view post-training through three interacting paradigms: fine-tuning (improving task and domain performance), alignment (shaping behavior to satisfy preferences and safety constraints), and reasoning enhancement (strengthening multi-step inference and planning). Within this space, three families of efficiency methods are especially consequential for reducing the cost of adaptation: parameter-efficient fine-tuning (PEFT), which limits the trainable subset of parameters or adds lightweight modules; model compression, which reduces memory and compute demands via quantization, pruning, or low-rank approximations and must often be co-optimized with adaptation; and knowledge distillation, which transfers expensive post-training outcomes to smaller or cheaper models for deployment. This thesis advances efficiency within and across these families, emphasizing that efficient post-training must be assessed under realistic constraints: repeated adaptation, privacy limitations, distributed computation, and interactive agent environments.The thesis presents four complementary contributions, ShareLoRA, AMAQ, LinguaLinked, and CCPO, each addressing a distinct bottleneck in the post-training lifecycle. First, ShareLoRA targets parameter- and storage-efficiency in PEFT by reducing redundancy in low-rank adaptation. Standard LoRA improves adaptation cost by training low-rank updates, but repeated specialization across many domains still accumulates trainable parameters and stored adapters. ShareLoRA addresses this repeated-adaptation setting by sharing low-rank matrices across layers while preserving sufficient flexibility for domain transfer. This design reduces the trainable parameter count relative to standard LoRA (reported reductions ranging from 44% to 96% in trainable parameters) while maintaining competitive performance across classification and generation tasks, and it supports robust behavior under continual fine-tuning where models must be updated multiple times over evolving data streams. As a result, ShareLoRA advances a practical route to scalable specialization: lowering both the cost to adapt and the footprint of many task-specific variants.Second, AMAQ targets communication-efficiency in collaborative post-training, where data is private or distributed and training must proceed under constrained bandwidth. In split-learning-style settings, transmitting intermediate activations and their gradients can dominate the end-to-end cost, and aggressive quantization can destabilize optimization, particularly at ultra-low bitwidths. AMAQ introduces an adaptive mixed-bit activation and gradient quantization strategy that progressively reduces precision (e.g., from 6--8 bits toward 3--4 bits) while allocating bit budgets based on layer-wise and feature-wise importance via bit regularization. This adaptive schedule mitigates representation collapse and improves training stability, enabling collaborative post-training to remain effective under tight communication budgets. In doing so, AMAQ demonstrates that quantization can be treated as a first-class component of post-training optimization rather than only an inference-time compression tool.Third, LinguaLinked addresses system-efficient deployment, focusing on privacy- and reliability-motivated scenarios where inference must run locally, but a single device cannot host a full model due to memory constraints. Rather than relying on centralized servers or reducing capability via aggressive downsizing, LinguaLinked enables decentralized inference by partitioning the model across multiple trusted mobile devices such that each device hosts only a fraction of the computation and parameters. The system combines optimized model assignment (segmenting the network and matching segments to device capabilities), optimized transmission (maintaining structured data flow between segments while preserving model structure), and a runtime load balancer that monitors performance and redistributes work to avoid bottlenecks. This line of work connects post-training to its downstream value: efficient adaptation is most impactful when specialized models can actually be served under realistic resource constraints without centralizing sensitive data.Finally, CCPO focuses on efficient agentic post-training for vision-grounded GUI agents, where reinforcement-learning-style optimization and long-horizon interaction drastically amplify iteration costs. In multimodal GUI/web navigation, agents must ground language in visual state, select precise actions, and succeed over long trajectories with sparse or delayed feedback. These properties exacerbate the dominant costs of the post-training loop: expensive on-policy rollouts, trajectory-level evaluation, and severe context inflation as interaction history grows. CCPO addresses these challenges by coupling policy optimization with coordinate-aware visual compression to preserve long-horizon context while reducing token and compute costs. It introduces Coordinate-Aware Spatial Compression (CASC), which aggregates coordinates across multiple rollouts to identify task-relevant regions and progressively narrow historical attention around key visual areas, and a distance-based advantage that provides a denser learning signal than binary correctness by rewarding progress in spatial grounding. In reported experiments, CCPO achieves strong performance across multiple GUI benchmarks while enabling up to 55% token compression and up to 3.8 × times training speedup, illustrating how capability and efficiency can be improved jointly in agentic post-training.Taken together, these contributions support a unified claim: efficient post-training requires coordinated optimization across how models are adapted, how training proceeds under practical constraints (communication, memory, privacy), how specialized models are deployed, and how agents are improved through interaction-heavy learning loops. Beyond the specific methods introduced, the thesis provides a set of reusable principles for efficient post-training: share structure to reduce redundant specialization state; allocate precision and bandwidth to where it matters most for learning stability; partition computation across trusted resources to satisfy memory and privacy constraints; and compress long-horizon context using task relevance while preserving the spatial and temporal information needed for robust decision-making. Collectively, this lifecycle view broadens the feasibility of customizing strong pretrained backbones and advances the practicality of LLM-based agentic systems under real-world resource constraints.

Main Content

This item is under embargo until March 11, 2027.