The people who use the Happy Pet Tech SaaS are groomers, kennel staff and vets. If a screen tells one of them "customer_id" is not allowed to be empty or E11000 duplicate key error, we have failed twice. They cannot act on that text, and we have shown a stranger the inside of our system.
In August and September we rebuilt how errors travel from the API to the screen. It has three parts.
1. The server decides what is safe to show
In Express, every error ends in one error handler. Ours sorts errors into two kinds.
Errors that our own code threw on purpose already have a message written for a person: “We couldn’t find that booking. It may have been removed.” These pass through.
Everything else is an error we did not plan for: a database error, a validation library’s raw text, a crashed dependency. For these, the handler checks if the message looks technical. If it does, the user gets plain copy, and the original text is kept on a separate field that only goes to the logs.
// simplified
function errorConverter(err, req, res, next) {
if (!(err instanceof ApiError) && looksTechnical(err.message)) {
const friendly = new ApiError(statusFor(err), GENERIC_MESSAGE);
friendly.debugMessage = err.message; // for logs, never for the response
return next(friendly);
}
next(err);
}
looksTechnical is a list of 14 patterns. Some examples of what it catches:
- validation phrasing:
"field" is required,must be a valid,must be a string - field names in
snake_caseorcamelCase - database words:
ObjectId,duplicate key,CastError - raw IDs: a 24-character hex string, a UUID
- network codes:
ECONNREFUSED,ETIMEDOUT - library names:
joi,axios,redis,jwt - the words
null,undefined,NaN - bare status text like
Internal Server Error - anything that looks like a stack trace or a file path
Validation errors get their own treatment. We validate request bodies with Joi, and Joi’s messages are written for developers. A small function turns them into plain sentences before they leave the server.
That function had a hole for a while. For Joi error types it did not know, it returned nothing, and the raw message went through. One of the types it did not know was the one for booleans, which our schemas use 328 times. The fallback now is the generic message, never the raw one.
2. The browser checks again
The server filter should be enough. We still run the same 14 patterns in the front end.
There are two reasons. Some messages never pass through our server: a network failure, or an error string that comes out of a library in the browser. And during a deploy, a new front end can talk to an older API for a short time.
export function pickErrorMessage(err, fallback = GENERIC_RETRY) {
const msg = err?.response?.data?.message ?? err?.data?.message ?? err?.message ?? null;
if (typeof msg !== 'string' || !msg.trim()) return fallback;
if (TECHNICAL_PATTERNS.some((re) => re.test(msg))) return fallback;
return msg;
}
If the message is empty or looks technical, the user sees “Something went wrong. Please try again.”
Some failed requests should show nothing at all:
- A 401 means the session ended. The app goes to the login page. A red box on top of that helps nobody.
- A cancelled request is not an error. The user left the page and we stopped the call.
- A suspended account has its own full screen.
3. Every response has a reference
The generic message is safe, but it tells us nothing either. So every response from the API carries a short reference in a header, like HP-CNAK7V1B.
const REFERENCE_ALPHABET = '0123456789ABCDEFGHJKMNPQRSTVWXYZ';
const generateReference = () => {
const bytes = crypto.randomBytes(8);
let code = '';
for (let i = 0; i < bytes.length; i++) {
code += REFERENCE_ALPHABET[bytes[i] % REFERENCE_ALPHABET.length];
}
return `HP-${code}`;
};
Look at the alphabet. It has no I, L, O or U. This is Crockford’s base32. It leaves out the letters that people confuse with digits when they read a code aloud. A customer on the phone with support can read HP-CNAK7V1B without saying “is that a zero or an O”.
The server always makes this code itself. If a client sends its own request ID, we keep it separately and do not reuse it. A reference that a client can choose is not a reference we can trust in the logs.
When an error is worth keeping, it is saved to an error log with that reference, the real message and the stack. Not every error qualifies. All 5xx errors are saved. Most 4xx errors are saved too, except the ones that are normal traffic: a 404 for a route that does not exist, a rate limit, an expired token.
Two limits protect the log from a bad day. The same error is saved at most 20 times a minute, and the whole process saves at most 300 rows a minute. When one broken endpoint fails 5,000 times, we need to know it is failing. We do not need 5,000 copies.
One CORS detail is easy to miss. A browser does not let JavaScript read a custom response header from another origin unless the API lists it in Access-Control-Expose-Headers. If your front end always sees undefined for the request ID, this is why.
What the browser records
When something fails in the browser, the front end reports it to the API with a short trail of what happened just before: the last 20 steps. A step is a page change, a click, or an API call with its method, path, status, duration and reference.
Two rules keep customer data out of that trail. The path is stored without its query string. And the report is only sent when there is a real crash, at most 10 per minute from one tab, and the same error only once a minute.
With the trail and the reference, most bug reports can be followed without asking the customer what they clicked.
What I would copy from this
- Filter on the server first. The browser filter is a second net.
- Keep the original message for the logs. A filter that throws it away makes debugging impossible.
- Give every response a short code that a person can read aloud.
- Put a limit on how many times the same error is stored.