TraceYield

How to Assess AI-Assisted Programming Assignments in HBO-ICT

A practical framework for assessing programming work when students legitimately use generative AI or coding agents, without reducing assessment to detection or manual authorship.

TY

TraceYield Insights

13 min read · Updated August 10, 2026

TraceYield

AI engineering evidence

How to Assess AI-Assisted Programming Assignments in HBO-ICT

The assessment problem is larger than authorship

A student submits a working application. The code is readable, the tests pass, and the interface behaves as required. A coding agent also contributed substantially during the work. The lecturer now has to assess more than whether the final artifact runs: what did the student understand, decide, verify, and learn while producing it?

That is not the same question as “was AI used?” In a professional programming context, tool use may be legitimate and expected. The educational question is whether the assessment still provides valid evidence for the learning outcomes. The answer requires a deliberate combination of final code, explanation, decisions, verification, and selected evidence from the development process.

Begin with the learning outcomes

List the outcomes the assignment is meant to assess before deciding what AI evidence to collect. If the outcome is implementing a tested service, the assessment needs behavior, design, tests, and the student’s ability to explain trade-offs. If the outcome is debugging, the route from symptom to diagnosis matters more than the number of lines changed. If the outcome is independent programming fundamentals, some controlled work without AI may be necessary.

This is constructive alignment applied to an AI-assisted setting. The permitted tool use, teaching activities, assignment, and evidence should point at the same outcome. A blanket ban can be misaligned when AI literacy is part of the professional goal; unrestricted take-home work can be misaligned when the intended outcome is unaided fluency.

Use several evidence layers

Final artifact evidence includes the running program, source code, tests, documentation, and deployment or demonstration. Process evidence can include the initial framing, approaches explored, meaningful changes, debugging steps, rejected alternatives, verification, and a short reflection. Explanation evidence can come from a code walkthrough, targeted questions, or a live modification of the submitted work.

The layers should be proportionate. A lecturer does not need a complete transcript of every prompt or every keystroke. The goal is a defensible sample of how the student worked and whether the student can connect the delivered result to the decisions behind it.

A practical assessment structure

The explanation should be targeted rather than a ritual defense of every line. Ask the student to explain one architectural choice, diagnose one failure, modify one behavior, and identify one limitation of the delivered solution. These prompts produce evidence of understanding without pretending that a conversation reconstructs the entire history perfectly.

What to assess in the process

Useful criteria include problem framing, choice of approach, quality of context given to the tool, evaluation of suggestions, debugging, testing, independent modification, and reflection on limitations. “Used AI effectively” is too vague to grade. Describe observable evidence instead: the student rejects an unsuitable approach for a stated reason, verifies a generated claim, or adapts code when the first solution fails.

Do not turn those observations into an ownership or AI-competence score. A rubric can describe levels of evidence for a learning outcome, but it should not claim to measure a person’s character or assign a universal independence percentage.

Handle legitimate differences in process

Programming tasks do not all produce the same trajectory. A familiar task may converge quickly. A new domain may require broad exploration. A student may use an agent to explain an error, generate tests, or compare designs rather than write the final implementation. The rubric should reward justified decisions and verification, not a preferred number of prompts or revisions.

When evidence is ambiguous, ask a focused follow-up question or request a small transfer task. Avoid treating uncertainty as proof of misconduct. Assessment integrity is strengthened by better evidence and fair procedure, not by confident guesses about authorship.

A compact rubric checklist

Before release, ask: are the learning outcomes explicit; is AI use described in student-facing language; does the assignment require meaningful decisions; are process and explanation evidence feasible at the class size; are the verification expectations clear; is there a fair way to handle uncertainty; and does the workload fit the credit and staffing available?

After the first run, review where the evidence helped and where it created bureaucracy. Programming assessment should remain a learning activity, not become a data-collection exercise. The design is successful when it gives the lecturer a better view of the intended learning while giving students a clear route to demonstrate it.

The TraceYield perspective

The final code shows what was delivered. A trajectory view can add context about how the work moved from task to result: where the student explored, introduced context, changed direction, verified, and made decisions. That context can support a lecturer’s judgment, but it does not replace the rubric, conversation, or professional assessment of the work.

Example: assessing a small service change

Suppose students extend a service that validates customer records. The final code and tests are assessed for behavior. The process evidence asks the student to identify one requirement that was ambiguous, explain one design choice, show how an agent suggestion was evaluated, and demonstrate how the solution behaves for an edge case. The evidence is compact but connected to the learning outcome.

A lecturer can then ask a different question of a small sample: “What would you change if validation had to become asynchronous?” The question is not intended to catch students. It tests whether they can transfer the relevant design idea beyond the exact generated patch.

Moderate before scaling

Before using a new rubric across a cohort, two lecturers should assess a small sample and compare what they inferred from the evidence. If one assessor rewards a long transcript and another rewards a concise explanation, the design needs clearer criteria. Moderation should focus on the learning outcome, not on agreeing that every student followed the same path.

This small investment can prevent a much larger workload later. It also gives students a more predictable assessment experience.

Use the assignment to create evidence

An assessment becomes easier to interpret when the task naturally produces the evidence the lecturer needs. A vague “build an application” prompt creates a final artifact but leaves the route under-specified. A better brief names a user or stakeholder, includes a constraint that requires judgment, asks for a testable behavior, and creates one point where the student must explain or revise a choice.

For example, a student might receive an existing codebase with an incomplete requirement and a known operational constraint. The student can use an approved agent, but must submit a short framing note, a design decision, tests for a boundary case, and a reflection on one generated suggestion that was changed or rejected. The evidence is not an audit trail of everything. It is a designed part of the learning activity.

Separate grading from investigation

A lecturer should not need to investigate every submission at the same depth. The rubric can identify the evidence everyone provides, while a sample or risk-based follow-up provides more direct explanation where the work is ambiguous. This keeps assessment manageable and avoids treating every student as a suspected integrity case.

The follow-up can be a short code walkthrough, a related debugging task, or a question about a design trade-off. It should be announced as part of the normal assessment model. Students are more likely to see it as a learning opportunity when it is not introduced only after a detector or intuition raises suspicion.

Questions for the marking team

Can two assessors identify the same learning evidence? Can a student demonstrate the outcome without exposing private tool data? Does the rubric reward verification and judgment rather than activity volume? Does the evidence fit the time available? These questions should be answered before the assignment is published, not after grading begins.

A short moderation exercise with anonymized samples can expose ambiguous language early. It is also a chance to check that the assessment remains accessible to students who use different tools or need reasonable accommodations.

Agent completion does not always mean engineering completion.

TraceYield engineering note

References

Pilot Program

Understand the WHY behind your engineering AI usage.

TraceYield evaluates trajectory evidence instead of stopping at spend totals. Join the private pilot to review AI coding usage with engineering context, security controls, and developer trust.

Join the TraceYield private pilot