There is a shortcut in every test automation tool, and it is always the same one: let the test step run a command. Give the engine a shell and it can do anything the system can do, so it can test anything. Nobody has to write a new step type ever again.
It is also the design that loses you the argument in an inspection, and it is worth understanding exactly why before you accept it in a tool you buy or build.
What a shell in the engine actually costs you
Three things, in rising order of seriousness.
- The test becomes unreviewable. An approver looking at a protocol step that reads run: ./verify.sh --env prod has approved a filename, not a test. What the script does at run time is not what they signed.
- The test can change without change control. The protocol is approved and frozen; the script it calls is a file on a disk that anybody with access can edit. Your approved protocol now executes something nobody approved.
- The engine can do things a test should never do. A test that can run commands can write to the system under test, clean up after itself, or quietly fix the thing it was supposed to be checking. Once that is possible, every passing result carries an asterisk.
None of these are hypothetical failure modes of a badly-behaved team. They are properties of the design. A control you have to trust people not to use is not a control.
The alternative: a closed vocabulary
The other approach is to decide, in advance, the complete list of things a test step is allowed to be — and to make anything outside that list impossible rather than discouraged.
A workable vocabulary for GxP protocols is smaller than people expect:
- Navigate to a named screen or endpoint.
- Enter a value into a named field.
- Click a named control.
- Assert that a named thing is present, absent, equal to, or within a range.
- Check a file exists, has a given checksum, or has given permissions.
- Check a service is running, or reports a given version.
- Capture a screenshot or a named log region as evidence.
Every one of those is fully described by its parameters. An approver reading the step knows precisely what will happen, because there is nothing else it can do. That is the entire point.
What you give up, honestly
You give up the long tail. Roughly one test in ten wants something the vocabulary does not have, and you now have two options rather than one: extend the vocabulary through your own change control, or execute that test by hand.
Both are slower than writing a script. That is the trade, and it is worth stating plainly rather than pretending the closed design is free. What you get back is that the other nine in ten are reviewable by a quality person who does not read code, and that no approved protocol can change underneath you.
The vocabulary also grows in the right direction. When a step type is added deliberately, it gets specified, tested and documented once, and then every protocol can use it. A shell gives you a thousand bespoke scripts, each of which is somebody's private understanding of what the test does.
How to assess a tool on this
If you are buying rather than building, the question to ask a vendor is blunt: can a test step execute code that is not part of your product?
- If the answer is yes with no qualification, the tool is a scripting harness with a compliance skin. It can still be used, but the controls now have to live in your procedures and your access model, and you should plan the validation accordingly.
- If the answer is yes but only via a plugin the vendor ships and qualifies, that is a reasonable middle. Ask which plugins, and what their qualification covers.
- If the answer is no, ask what happens when a protocol needs something the vocabulary lacks. A tool with no honest answer to that has simply moved the problem.
We took the closed route in GxP Copilot, with no general command step by design, and the constraint has been productive more often than it has been annoying — mostly because it forces the question "what is this step actually verifying?" at design time rather than at review.
Where to go next
Explore GxP Copilot for AI-native validation, TraceDraft for source-traceable clinical documentation, or book a demo to see either on your own data.
