Evaluating Beyond the Screen: Collective Assessment of AI-Generated Business Plans with Resource-Constrained Entrepreneurs

2026-08-17Human-Computer Interaction

Human-Computer Interaction
AI summary

The authors studied how entrepreneurs use AI tools like ChatGPT to create important business documents but found errors can cause problems. Instead of having people check AI output alone on a screen, they explored group evaluation methods. They enhanced a business-planning tool called BizChat to link AI-generated claims back to the users' original inputs and tested it in workshops. Their early results showed that group discussions and visual aids helped entrepreneurs better understand and check the AI's work using peers' knowledge and printed materials.

generative AIChatGPTentrepreneurshipbusiness planningAI evaluationthink-pair-shareuser interfacedigital literacycollaborative learning
Authors
Qi Zhao, Marjory Pineda, Ketul Chhaya, Aakash Gautam, Yasmine Kotturi
Abstract
Entrepreneurs increasingly use end-user generative AI technologies such as ChatGPT for high-stakes documents like loan applications and business plans, where AI-generated errors---a wrong price, a fabricated product---can affect loan or funding outcomes. Current approaches to supporting evaluation of AI-generated text assume a single user assessing output alone, on screen. This can be especially demanding for resource-constrained entrepreneurs, whose digital and AI skills vary widely. In this early-stage work, we explore how evaluation might instead be organized in a group setting and completed as a collective activity. We extended BizChat, an AI-powered business-planning tool, with an evaluation module that links each generated claim to the entrepreneur's original input. We partner with community organizations in Maryland---embedding BizChat within various entrepreneurship programs---where workshop attendees (N=14) evaluated their plans through think-pair-share discussion. Early findings suggest interface scaffolds like claim-to-input links primed attendees with concrete, personal evaluations, which the group setting then extended beyond the screen: attendees requested printed copies, used rubrics to compare across plans, and drew on peers' knowledge to verify what they could not easily judge alone.