MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations
This paper evaluates how well large language model (LLM) agents can perform long-term tasks in a real-world e-commerce setting, where they need to make decisions over time and adapt to changing conditions. Practitioners in e-commerce and AI research can learn from this study to improve the performance of LLM agents in similar environments.