Preprint
Concept Paper

This version is not peer-reviewed.

General Education in the Age of Artificial Intelligence: Assessing Student Judgment Beyond the Final Product

Submitted:

18 August 2026

Posted:

20 August 2026

You are already at the latest version

Abstract
Generative AI challenges the evidence that universities use to certify general education learning. A polished final product can no longer show, by itself, that a student can reason, write, evaluate information, use data, or make responsible judgments. This paper argues that general education should move from course completion alone to a competency-and-practice model. In this model, students first demonstrate core abilities without AI assistance. They then use AI in structured assignments where they must disclose use, verify outputs, revise results, and explain their decisions. The paper applies this model to thinking, communication, quantitative and data reasoning, AI and information literacy, ethical and civic judgment, and integrative learning. It also proposes a dual-condition assessment design that separates independent competence from documented AI-assisted performance.
Keywords: 
;  ;  ;  ;  ;  ;  ;  
Subject: 
Social Sciences  -   Education

1. Introduction

General education is the institution’s promise that all students will learn core skills. Students in every major must complete part of their degree in general education, often 30 to 45 of 120 credits. These courses aim to expose students to multiple fields, build key literacies, and prepare them for civic life. Generative AI challenges this promise. Students can now generate clear writing, cross-disciplinary claims, calculations, and ethical positions in seconds. The task shifts from writing a response on their own to judging one. Students must question assumptions, check claims against evidence, place ideas in a discipline, and take responsibility for how they use the work. This matters most in general education because it is where universities claim they certify the same skills that AI can mimic.
In the absence of a deliberate institutional response, one potential consequence is credential inflation, as institutions may find it increasingly difficult to demonstrate that awarded credits reflect independently demonstrated competencies. Students accumulate general education credits while AI mediates the intellectual work those credits are designed to represent. The appropriate response is neither prohibition nor unrestricted delegation, but a redesigned program where for many foundational general education competencies, institutions may choose to establish independent competency before extensive AI-supported work.
This paper presents a conceptual model and institutional design proposal informed by emerging evidence, rather than a validated curricular framework.

2. What General Education Is and Why It Matters

The academic literature defines general education more broadly than state law does. General education develops four things in all students: breadth of knowledge, intellectual and practical skills, personal and social responsibility, and integrative learning across disciplines, regardless of major (Association of American Colleges and Universities 2007). These outcomes are operationalized through the AAC&U VALUE rubrics, which provide assessment instruments for each domain (Association of American Colleges and Universities 2009). The Boyer Commission (1998) called for inquiry-based undergraduate education along similar lines. Pascarella and Terenzini (2005) reviewed several hundred empirical studies and found that general education produces measurable gains in critical thinking, written communication, and civic orientation. Critically, those gains depend on pedagogy: courses that require active, cumulative engagement with a competency produce stronger outcomes than courses that rely on content exposure alone. That finding directly motivates the redesign proposed in this paper.
State law operationalizes these commitments through credit-hour requirements and subject-area mandates. Florida defines general education as required coursework in communication, mathematics, social sciences, humanities, and natural sciences, with core courses that develop academic and critical-thinking skills. Texas requires a 42-credit lower-division core covering knowledge of human cultures, personal and social responsibility, and practical skills including critical thinking, communication, and quantitative reasoning. Alabama uses a statewide general studies core organized around written composition, humanities, natural sciences, mathematics, and social sciences, with a 41- to 42-credit lower-division structure (Florida Legislature 2025,Texas Education Code 2025,Texas Administrative Code 2025,Alabama Transfers 2026). The statutory form varies across states. The underlying outcome commitments are consistent across states and correspond to the AAC&U framework.

3. Generative AI Challenges in GenEd Courses

Generative AI challenges general education differently from how it challenges discipline-specific or professional education. In a field, AI disrupts workflow. In general education, it challenges the rationale for certification. General education is the locus at which universities assert that thinking, written communication, and civic reasoning are competencies that can be taught, assessed, and credentialed. When AI produces grammatically proficient prose, structured cross-disciplinary arguments, and formally coherent ethical positions, the institutional question shifts from whether individual assignments require redesign to whether the credential retains the validity it claims.
Farazouli et al. (2024) provide direct empirical evidence for this concern. In philosophy, law, and education—disciplines relevant to many general education programs—university faculty could not reliably distinguish student-authored from AI-generated examination responses, and AI-generated responses received passing grades in a substantial share of cases. This result demonstrates that final written products, the primary assessment vehicle in many general education settings, no longer constitute sufficient evidence of the competencies they are intended to certify. The institutional response cannot be a ban on using AI. AI is embedded in the professional and civic environments students will inhabit. It requires a redefinition and operationalization of what general education certifies that AI cannot substitute: disciplinary judgment, contextual reasoning, audience-responsive communication, and accountability for the consequences of the conclusions reached.
Controlled experimental evidence on AI-supported learning is concentrated in mathematics, introductory programming, and introductory physics—domains that constitute a small fraction of most general education programs. Three randomized studies define the methodological state of the field. Bastani et al. (2025) found that unrestricted access to GPT-4 in a high school mathematics trial improved assisted-practice performance but reduced subsequent unassisted examination performance; a pedagogically constrained, hint-based tutor preserved learning gains. Bassner et al. (2026) found that students who used ChatGPT or a structured AI tutor performed better on programming exercises, but did not show better performance on knowledge tests or code-comprehension measures. Kestin et al. (2025) found that a carefully designed AI tutor in physics improved immediate learning compared with an active-learning class.
The key distinction is that AI improves learning when carefully scaffolded and grounded in instructor-designed solutions, rather than used as an unrestricted answer generator. AI-mediated instruction can inflate task performance while concealing deficits in independent competence. Of greater significance for general education is the disciplinary gap in the evidence base. Writing instruction, ethical reasoning, civic argument, and cross-disciplinary synthesis—the domains in which general education’s distinctive purposes are most fully realized—have been the subject of few controlled AI learning trials. Curriculum designs proposed in this paper are evidence-informed proposals derived from established pedagogical principles, not findings validated by experimental investigation.
Table 1. Domains, supported findings, and suggestive actions for GenEd curriculum.
Table 1. Domains, supported findings, and suggestive actions for GenEd curriculum.
Domain What Research Directly Supports Proposed Responses for GenEd Curriculum
Critical-thinking pedagogy Explicit, active, content-embedded instruction produced measurable improvement in assessed thinking in a required communication course (Mazer et al. 2008). Inquiry must be named as an outcome, taught within disciplinary content, and assessed directly.
Quantitative and STEM AI tutoring Unrestricted AI access can reduce subsequent independent performance; pedagogically constrained, hint-based systems preserve or improve learning (Kestin et al. 2025). Deploy hint-first, instructor-grounded AI for quantitative practice; retain independently assessed competency checks.
Writing assessment in humanities and social sciences AI-generated examination responses in philosophy, law, and education received passing grades and were not reliably identified from final prose alone (Farazouli et al. 2024). Assess writing competency through process evidence, documented source use, and oral examination—not final written products alone.
AI literacy and ethics AI literacy comprises understanding, use, evaluation, and ethical governance of AI systems; model hallucination, privacy risks, inequity, and opacity are documented harms (Ng et al. 2021,Yan et al. 2024). Embed evaluation, disclosure requirements, data stewardship, and viewpoint diversity as explicit general education outcomes.
Writing and cognitive engagement A small EEG study reports an association between LLM-assisted drafting and reduced neural engagement and lower authorial ownership (Kosmyna et al. 2025). Require independent drafting prior to AI-assisted critique on design grounds; this evidence is insufficient to support the prescription independently.

4. Redesigning the General Education Curriculum

4.1. Adopt a Competency-and-Practice Model

General education should not rely only on requiring students to take courses in different subject areas. That model assumes that completing courses in writing, mathematics, humanities, social sciences, and natural sciences is enough to demonstrate broad undergraduate learning. In the AI era, this assumption is weaker. A student may complete work in each area with substantial AI assistance while showing limited independent ability to write, reason, interpret evidence, analyze data, or make ethical judgments.
A competency-and-practice model addresses this problem only if competency checks are designed to directly verify students’ thinking. General education competencies should therefore be assessed under two conditions. First, students must demonstrate core ability without AI assistance through in-class writing, oral explanations, short proctored tasks, handwritten or monitored problem-solving, or live defense of their reasoning. Second, students may use AI in structured assignments, but they must document how the tool was used, verify its claims, revise its output, and explain what they accepted or rejected. This approach is consistent with evidence that explicit, embedded, and cumulative instruction can improve assessed critical-thinking performance in general education settings (Mazer et al. 2008), and with AI literacy frameworks that emphasize understanding, use, evaluation, and ethical governance of AI systems (Ng et al. 2021,Yan et al. 2024).

4.2. AI as an Instructional Tool

The default student orientation toward AI in general education courses must be interrogative rather than delegative. In foundational courses, the cognitive operations of drafting, arguing, and revising constitute the mechanism through which the competency under assessment is developed. Transferring those operations to AI before competency acquisition undermines the course’s educational function. AI can serve five pedagogically defined roles that preserve student intellectual agency:
  • Socratic interlocutor: prompt students with targeted questions and partial information rather than complete solutions;
  • Adversarial critic: identify unstated assumptions, logical gaps, missing evidence, or alternative interpretations in the student’s own argument;
  • Generative practice partner: produce low-stakes examples, datasets, counterexamples, and rehearsal prompts for course preparation;
  • Calibration instrument: provide a first-pass response that students evaluate against a rubric, instructor feedback, and peer review; and
  • Simulation agent: adopt the perspective of a historical actor, policy stakeholder, civic audience, or disciplinary interlocutor relevant to the topic under study.
These roles preserve the student as the primary intellectual agent and instantiate the form of AI engagement most relevant to professional and civic contexts: governed, disclosed, and critically evaluated use (Yan et al. 2024).
Figure 1. The Competency-and-Practice Model for General Education in the AI Era.
Figure 1. The Competency-and-Practice Model for General Education in the AI Era.
Preprints 229003 g001

5. Teaching Core General Education Capabilities

5.1. Critical Thinking

A student who cannot identify an unsubstantiated AI claim, a fabricated citation, or a biased framing does not possess critical thinking at the level a general education credential should certify. In every general education course, students should systematically:
  • decompose an AI-generated response into constituent claims, supporting evidence, embedded assumptions, and logical inferences;
  • verify factual claims against assigned readings, primary sources, and authoritative reference databases;
  • compare outputs across alternative prompts or model configurations for systematic differences in framing, emphasis, and omission; and
  • revise the final argument under their own authorship, with explicit documentation of which AI contributions were incorporated, modified, or rejected.
The reasoning memo—a structured artifact in which students record their initial position, the nature of AI contributions, the evidence used to evaluate those contributions, and the justification for final revisions—renders visible the analytic and reflective operations that general education is designed to develop (Mazer et al. 2008). The prescribed assignment sequence is: (1) independent analysis; (2) AI critique of that analysis; (3) source-based verification of the critique; and (4) final revised argument accompanied by the reasoning memo. This sequence ensures that competency development precedes AI-augmented practice.

5.2. Communication

Generative AI can make student writing appear more polished—improving grammar, organization, and surface clarity. For this reason, a final paper alone is no longer sufficient evidence that a student has strong communication skills. Communication competency should be assessed through process-based evidence, not only through final written submissions:
  • staged drafts and source logs that document the intellectual development of the student’s argument;
  • oral examination or brief viva on the principal claims of submitted written work;
  • in-class or proctored writing to establish an independently verified baseline of fluency; and
  • audience-specific oral presentations and structured deliberative exercises.
These instruments institutionalize the professional norm that general education should certify: AI assistance is permissible, but human authorial accountability for accuracy, source integrity, and consequences is not transferable.

5.3. Quantitative and Data Reasoning

This paper introduces the EMAI framework for quantitative reasoning in general education: Estimate, Model, Audit, Interpret. The framework draws on quantitative literacy education (Steen 2001) and model-based reasoning (Niss et al. 2007). It extends both to a specific problem: students must interrogate AI-generated outputs, not simply accept them. The four steps are:
  • Estimate: specify the expected magnitude and direction of the result prior to AI engagement;
  • Model: direct AI to propose a calculation method, code implementation, or statistical interpretation;
  • Audit: verify assumptions, variable definitions, units, and numerical results against an independent calculation; and
  • Interpret: articulate precisely what the result supports and what inferential claims it does not warrant.
Data literacy in this framework is connected to ethical and responsible use. Students must evaluate where data come from, how they were collected and measured, whether they are complete and appropriate for the question, and what privacy or practical consequences may follow from data-based decisions (Yan et al. 2024).

5.4. AI and Information Literacy

AI literacy should be part of general education, not only a topic for technical courses. Students need to understand what AI systems can do, use them appropriately, check their outputs, and recognize the ethical issues that may arise (Ng et al. 2021). These skills should be taught within existing general education courses: ethics courses examine responsible AI use and the limits of AI-generated ethical reasoning; statistics courses teach output verification and uncertainty; writing courses address disclosure, source verification, and authorial accountability. Core content should encompass:
  • Capabilities and limitations: the mechanisms by which language models produce plausible but potentially inaccurate outputs, and the domains in which errors are systematic;
  • Output verification: procedures for checking factual claims, source citations, and quantitative results before academic or professional use;
  • Data stewardship: criteria for determining what information should not be submitted to third-party AI systems;
  • Appropriate access and use: differences in tool availability, data quality, language support, and user preparation that affect how students evaluate AI outputs; and
  • Professional and disciplinary judgment: application of field-specific standards to consequential decisions informed by AI outputs.

5.5. Ethical and Civic Judgment

Generative AI creates new issues that general education ethics courses must address: academic integrity, data privacy, responsible AI use, and the role of AI in civic and professional decision-making. Course policies should state clearly how students may use AI on each assignment. A simple taxonomy can distinguish four cases: AI use is not allowed; AI is allowed only for practice; AI-assisted work is allowed with disclosure; or AI-collaborative work is allowed with a fully documented process—each tied to the assignment’s learning goal.
Integrity policy should not rely primarily on AI detection software. Farazouli et al. (2024) found that faculty could not reliably distinguish AI-generated text from student-authored text. Detection software output constitutes grounds for further review, not definitive proof of misconduct. Clear use rules and process-based evidence provide a stronger basis for evaluating student work.

5.6. Integrative and Applied Learning

Integrative learning means applying knowledge from different fields to a specific problem, using evidence, understanding the setting, and explaining the limits of a recommendation. AI can summarize ideas from different fields, but it cannot take responsibility for a recommendation or judge a local problem in the way a student must. Integrative general education assignments should use local cases, institutional or community data, and oral discussion. Students should use AI as a tool—checking its output, adapting it to the case, and explaining why they accepted or rejected its suggestions. This gives institutions direct evidence that students can apply judgment, not only produce AI-assisted text.

6. Assessment Design for General Education in the AI Era

The foundational assessment question for general education is no longer whether students can produce a response, but what students can demonstrate independently and what they can accomplish with documented, governed AI assistance. A valid general education program must assess both conditions with direct evidence mapped to stated learning outcomes.
Table 2. General education assessment purposes and corresponding design responses.
Table 2. General education assessment purposes and corresponding design responses.
Assessment Purpose Design Response
Establish foundational competency Administer in-class, oral, or proctored assessments requiring students to demonstrate and explain reasoning without AI assistance.
Develop AI-augmented practice Permit disclosed AI use in drafting, feedback, and iterative revision; evaluate the quality of verification, revision, and reflective documentation—not the surface characteristics of the final product.
Assess integrative judgment Assign locally grounded cases with specified stakeholder constraints and ethical dimensions; require oral defense and disciplinary artifacts that cannot be produced by AI without domain-specific contextual knowledge.
Maintain integrity and equity Specify permitted AI use prior to each assessment; provide institutionally administered, privacy-compliant tool access; adjudicate suspected violations through process evidence rather than detection software output (Farazouli et al. 2024,Yan et al. 2024).
Satisfy accreditation requirements Map direct assessment evidence from both independent and AI-augmented conditions to general education learning outcomes in the institutional assessment matrix; document the improvement feedback loop for each outcome.
A prohibition on AI use is neither enforceable nor pedagogically coherent: authorship cannot be determined reliably from final text alone, and AI literacy is itself a general education outcome. Unrestricted AI delegation is equally indefensible—AI-supported task performance can conceal deficiencies in independent competence that a general education credential should not certify. The dual-condition model assesses students in two ways: (1) through restricted, in-class, or oral assessments of independent competence; and (2) through process records, reasoning memos, and documented verification of AI-augmented work.

7. Implementation Challenges

General education serves many students across the university and relies heavily on instructors who teach multiple sections or large courses. An assessment method that works in a 25-student seminar may not work in a 400-student course across 16 sections. Three structural design choices make the model operationally sustainable:
  • Stratified sampling rather than universal administration: Institutions need not apply oral verification universally. Sampling approaches similar to accreditation assessment methods may provide sufficient evidence of competency at the program level.
  • Substitution rather than addition: replace some final-product grading with process-based checkpoints—draft reviews, source checks, reasoning memos, or short oral explanations—keeping grading workload comparable while producing better evidence of student learning.
  • Centralized instrument development: produce common rubrics, model assignments, disclosure templates, and calibration protocols once at the general education program level and disseminate across all sections.
With respect to access equity: where institutionally administered AI access cannot be guaranteed, assessments must revert to AI-restricted conditions. Assessing AI-augmented competency without equitable tool access creates a grading structure in which students are evaluated in part on the quality of commercially available tools they can privately afford—inconsistent with the equity obligations of a publicly accountable institution (Yan et al. 2024).

8. Limitations and Conclusions

The controlled experimental evidence supporting the instructional designs in this paper is concentrated in STEM domains that represent a minority of most general education programs. Whether pedagogically constrained AI tutoring produces durable competency gains in writing-intensive, interpretive, or civic-reasoning general education courses has not been established through randomized trials. Recommendations for thinking, ethical reasoning, oral communication, and integrative learning come from established teaching research and curriculum design principles, not from AI-specific experiments in those areas. AI creates different assessment challenges across general education domains, and therefore implementation strategies may vary by discipline.
Generative AI changes the evidence needed to support general education. A polished answer is no longer enough. Students must demonstrate how they reason, how they verify information, and how they take responsibility for conclusions. General education should therefore assess both independent work and documented AI-assisted work.
The redesign proposed in this paper gives institutions a practical way to do this. It defines what students must demonstrate without AI and what they may do with AI support. The goal is direct: general education should certify student judgment, not only completed coursework.

References

  1. Association of American Colleges and Universities. 2007. College Learning for the New Global Century. Washington, DC: Association of American Colleges and Universities. [Google Scholar]
  2. Association of American Colleges and Universities. VALUE: Valid Assessment of Learning in Undergraduate Education, 2009. Accessed. (accessed on August 2026).
  3. Boyer Commission on Educating Undergraduates in the Research University. 1998. Reinventing Undergraduate Education: A Blueprint for America’s Research Universities. Technical report. Stony Brook, NY: State University of New York at Stony Brook. [Google Scholar]
  4. Pascarella, E.T., and P.T. Terenzini. 2005. How College Affects Students: A Third Decade of Research. San Francisco, CA: Jossey-Bass, Vol. 2. [Google Scholar]
  5. Florida Legislature. General Education Courses; Common Prerequisites; Other Degree Requirements, § 1007.25. Florida Statutes, Title XLVIII, Chapter 1007 (2025 ed.), 2025. Amended through. pp. ch. 2024–101. (accessed on August 2026).
  6. Texas Education Code. § 61.821: Definitions. 2025, Accessed. (accessed on August 2026).
  7. Texas Administrative Code. 19 Tex. Admin. Code 4.28: Core curriculum. 2025, Accessed. (accessed on August 2026).
  8. Alabama Transfers. Areas I–V, 2026. Accessed. (accessed on August 2026).
  9. Farazouli, A., T. Cerratto Pargman, K. Bolander-Laksov, and C. McGrath. 2024. Hello GPT! Goodbye home examination? An exploratory study of AI chatbots’ impact on university teachers’ assessment practices. Assessment & Evaluation in Higher Education 49: 363–375. [Google Scholar] [CrossRef]
  10. Bastani, H., O. Bastani, A. Sungu, H. Ge, Ö. Kabakcı, and R. Mariman. 2025. Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences 122: e2422633122. [Google Scholar] [CrossRef]
  11. Bassner, P., B. Lenk-Ostendorf, R. Beinstingel, T. Wasner, and S. Krusche. 2026. Less stress, better scores, same learning: The dissociation of performance and learning in AI-supported programming education. Computers and Education: Artificial Intelligence 10: 100537. [Google Scholar] [CrossRef]
  12. Kestin, G., K. Miller, A. Klales, T. Milbourne, and G. Ponti. 2025. AI tutoring outperforms in-class active learning: An RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports 15: 17458. [Google Scholar] [CrossRef] [PubMed]
  13. Mazer, J.P., S.K. Hunt, and J.H. Kuznekoff. 2008. Revising general education: Assessing a critical thinking instructional model in the basic communication course. The Journal of General Education 56: 173–199. [Google Scholar] [CrossRef]
  14. Ng, D.T.K., J.K.L. Leung, S.K.W. Chu, and M.S. Qiao. 2021. AI literacy: Definition, teaching, evaluation and ethical issues. Proceedings of the Proceedings of the Association for Information Science and Technology Vol. 58: 504–509. [Google Scholar] [CrossRef]
  15. Yan, L., S. Greiff, Z. Teuber, and D. Gašević. 2024. Promises and challenges of generative artificial intelligence for human learning. Nature Human Behaviour 8: 1839–1850. [Google Scholar] [CrossRef] [PubMed]
  16. Kosmyna, N., E. Hauptmann, Y.T. Yuan, J. Situ, X.H. Liao, A.V. Beresnitzky, I. Braunstein, and P. Maes. 2025. Your brain on ChatGPT: Accumulation of cognitive debt when using an AI assistant for essay writing task. arXiv (preprint, not yet peer-reviewed). arXiv:2506.08872. [Google Scholar] [CrossRef]
  17. Steen, L.A. 2001. Mathematics and Democracy: The Case for Quantitative Literacy. Princeton, NJ: National Council on Education and the Disciplines. [Google Scholar]
  18. Niss, M., W. Blum, and P. Galbraith. 2007. Introduction. In Modelling and Applications in Mathematics Education. Edited by W. Blum, P.L. Galbraith, H.W. Henn and M. Niss. New York, NY: Springer, pp. 3–32. [Google Scholar] [CrossRef]
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.