- Main
Generative Models for Image-Based Plant Analysis: Thermal Image Super-Resolution and Plant Architecture Generation
- Yun, Heesup
- Advisor(s): Earles, Mason
Abstract
Vision models replace manual measurement of plants from an agricultural field, measuring plant traits with throughput beyond human capability and with greater consistency than manual methods. Recent vision models trained on web-scale datasets with billions of labeled examples have achieved near-human accuracy on natural image benchmarks. Yet, this scaling is unachievable for agricultural vision tasks where seasonal cropping cycles, sensor costs, and data annotation requirements restrict training data to orders of magnitude fewer samples. In breeding programs, for instance, evaluating thousands of genotypes across seasons forces a rigid trade-off between high-throughput, low-resolution remote sensing and low-throughput, high-resolution proximal measurement, typically yielding only thousands of images per season.To resolve this bottleneck, this dissertation develops generative-model-based sensing methods that increase effective data resolution and throughput without proportionally increasing field labor or sensor costs: (1) thermal image enhancement for low-cost thermal cameras via image alignment and super-resolution; (2) plant architecture generation via vision language models; and (3) plot-level simulation generation via vision language models with in-context learning.First, the resolution of low-cost thermal images was upscaled using pixel-level red–green–blue (RGB) and thermal alignment and a generative adversarial network (GAN)-based super-resolution algorithm. Thermal and RGB images were collected from a smartphone-mounted low-cost thermal camera. The collected RGB images were converted to pseudo-thermal images and aligned with low-resolution thermal images using template matching. Then the aligned RGB, pseudo-thermal, and low-resolution thermal images were fused and enhanced by the super-resolution algorithm, upscaling the resolution by a factor of four. Our results show an root mean squared error (RMSE) of 2.75 °C on temperature measurement and a structural similarity index measure (SSIM) of 0.63 on the test set, and demonstrate recovery of leaf edge details from low-resolution thermal images by combining shape information from the RGB images.Second, a plant architecture generation algorithm was developed using a vision language model with vision encoder and language decoder. Using a synthetic cowpea dataset generated from the Helios plant simulation library, we developed a tokenization method to serialize the nested structure of plant architecture and quantize continuous variables into discrete levels. Then, the vision language models were trained with 400,000 synthetic cowpea images and plant architecture pairs. The results show that our models can generate plant architecture from regular RGB images and outperform conventional regression-based methods in estimating leaf count and leaf area, with a mean absolute percentage error (MAPE) of 4.1% for leaf count and 3.2% for leaf area.The final chapter expands the plant-level architecture generation to plot-level simulation generation using open-weight vision-language models with in-context learning. The plant simulation configuration includes metadata (year, location, plant type, days after planting), plant count, plant locations, solar position, and chlorophyll content, which are needed for Helios-based plant simulator to render the simulated plot. The open-weight models were evaluated on their ability to generate plant simulation configurations in JavaScript Object Notation (JSON) format from an image, using five different in-context learning methods. Five in-context learning methods were designed to provide information gradually, from basic format-restriction instructions, to few-shot examples, to grounding hints. Results show that the vision language models can generate plant simulation configurations from an image, and the Gemma 4 31B model with few-shot examples achieves an average symmetric mean absolute percentage error (sMAPE) of 19.4%.