The integration of large-language-model (LLM) agents into materials science requires a balance between adaptive reasoning and reliable scientific execution. Here, we present a Harness framework that converts agent proposals into structured and validated actions, connects them to the Materials Project and deterministic machine-learning tools, and records settings and computational evidence. The present perovskite band-gap case study uses an expert-defined oxide/halide partition as an initial subclass strategy; it does not claim autonomous discovery of an optimal material taxonomy. Fourteen regressors are compared under pooled and subclass-specific training using identical test entries across 50 matched stratified splits. Subclass-specific modeling directly improves several linear, nearest-neighbor, and ensemble regressors, but the effect is model dependent and is negative for SVR and LightGBM. A direct training-only K-means baseline produces a different partition; the chemical rule performs better for several regressors, whereas K-means performs better for LightGBM. In the current implementation, Harness provides the reproducible execution pathway for evaluating this scientific hypothesis. More broadly, the framework is designed to compare and refine candidate subclasses, descriptors, models, and physical-validation tools using downstream performance feedback. This work therefore provides a foundational demonstration of adaptive, tool-integrated workflow construction for materials science.



