Compute and storagemedium3-5 years
Users upload invoices of up to 200 MB. The current endpoint accepts a multipart body in a Spring controller and streams it to S3 with the SDK; under load the service's heap and threads are exhausted. How would you redesign the upload, what does the service still own, and what does 'S3 is not a filesystem' mean for the rest of the design?
Stop proxying the bytes. The service authenticates the user, decides the object key, and signs a presigned URL that permits one PUT to that one key for a short time (fifteen minutes, say); the client uploads directly to S3 and the bytes never pass through the JVM. The service still owns everything that is a decision: who may upload, the key layout, the content type it signed for, and what happens after the upload, which an S3 event notification to SQS turns into an event instead of a poll. 'Not a filesystem' means there is no rename (it is copy then delete), no append, and directories are only a prefix convention, so keys are designed once, as invoices/<customer>/<date>/<uuid>, not moved around later.
PreviousA service on ECS reads an access key and secret from `application.properties` (injected from a CI variable) to talk to S3 and Secrets Manager. You are asked to 'harden' it. What should replace the key, how does the SDK find credentials with no configuration, and what does the policy on the new identity look like?Next An ECS service behind an ALB has its target group health check pointed at `/actuator/health`, which includes the database. During a 10-minute RDS incident the dashboard shows 0 healthy targets, yet some requests still reach the tasks, and ECS keeps replacing tasks that never finish starting. Explain the health check's semantics on an ALB, the fail-open behaviour, and what you would change.