SurveyUQBench: Recovering Complex-Survey Designs and Design-Based Uncertainty from Official Documentation

Main Article Content

Junxi Zhou

Keywords

benchmark evaluation, complex survey analysis, design-based uncertainty, large language model agents, survey design recovery

Abstract

Reliable public-survey analysis combines plausible coding and weighted point estimation with the recovery of the correct weights, strata, primary sampling units, replicate-weight convention, domain rule, degrees of freedom, and confidence-interval rule fro m official documentation. We introduce SurveyUQBench, a benchmark/resource that evaluates this complete path to design-based uncertainty. It contains 360 tasks drawn from 12 public survey programs and five variance mechanisms, together with official-document evidence qrels, evaluation-hidden design manifests, reference R analyses, and deterministic graders for evidence, design recovery, and estimate/standard-error/confidence-interval correctness. Reference analyses and clean replays cover the full resource, and mutation tests confirm that the deterministic grader distinguishes valid references from targeted mutations. Direct model responses achieve broad compliance with the response contract; strict NumericUQ and FullTask both record 0 under the joint criteria. On a fixed task-model subset, submission - native execution achieves 23 successful invocations; deterministic interface normalization increases this to 89, yielding 9 finite estimate/SE pairs and 2 NumericUQ passes —gold-input interventions further separate evidence, design, interface, and numerical stages. SurveyUQBench therefore provides a traceable evaluation resource, a reproducible baseline, and stage-specific signals for measuring future progress.

Abstract 8 | PDF Downloads 2

References

  • [1] Lumley, T. (2004). Analysis of complex survey samples. Journal of statistical software, 9, 1-19.
  • [2] Lumley, T. (2020). Survey: analysis of complex survey samples. R package version, 4(1).
  • [3] Li, Y., Zhang, Z., Ma, T., Wang, Z., Murugesan, K., Zhang, C., & Ye, Y. (2026). LongDA: Benchmarking LLM Agents for Long-Document Data Analysis. arXiv preprint arXiv:2601.02598.
  • [4] Kim, J., Jeong, D., Son, B., Kim, H., Kim, B., & Han, K. (2026, April). LAPS: Automating Hypothesis-CHI Conference on Human Factors in Computing Systems (pp. 1-19)
  • [5] Song, X., Lee, L., Xie, K., Liu, X., Deng, X., & Hong, Y. (2026). Statllm: A dataset for evaluating the performance of large language models in statistical analysis. Scientific Data.
  • [6] Zhu, Y., Ding, Y., Lai, P., Wang, L., Jing, B., & Chen, G. (2026). StatABench: Dataset and Framework for Evaluating Statistical Analysis Capabilities of LLMs. arXiv preprint arXiv:2606.22977.
  • [7] Lu, Y., Yang, R., Zhang, Y., Yu, S., Dai, R., Wang, Z.,... & Zhou, F. (2025). Stateval: A comprehensive benchmark for large language models in statistics. arXiv preprint arXiv:2510.09517.
  • [8] Ma, R., Shankar, S., Chen, R., Lin, Y., Zeighami, S., Ghosh, R.,... & Parameswaran, A. G. (2026). Can ai agents answer your data questions? a benchmark for data agents. arXiv preprint arXiv:2603.20576.
  • [9] Sun, Z., Zhong, S., Wen, D., Han, J., Li, G., Yan, Y.,... & Ruan, H. (2026). AgenticDataBench: A Comprehensive Benchmark for Data Agents. arXiv preprint arXiv:2607.01647.
  • [10] Zhu, Y., Jin, T., Pruksachatkun, Y., Zhang, A., Liu, S., Cui, S.,... & Kang, D. (2026). Establishing best practices in building rigorous agentic benchmarks. Advances in Neural Information Processing Systems, 38.
  • [11] Ghosh, A., Reuel, A., Chim, J., Kennedy, W. M., Yadav, S., Mickel, J.,... & Solaiman, I. (2026). Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting. arXiv preprint arXiv:2606.09809.
  • [12] Rothschild, D. M., Marla, J., Amaya, A., Barari, S., Buskirk, T. D., Cobb, C.,... & Webb, B. (2026). Responsible AI Integration in Survey Research.