Standard Site, cardyb and shortlinks

I’ve been doing some work around standard.site integration with the work blogs and I ran into something I found interesting. Shortlinks, especially when part of a post that was posted by a third party sheduler don’t show the integration on Bluesky (and similar clients).

Every enterprise publisher and open-source foundation (Google, Linux Foundation, CNCF, Mozilla) uses vanity shortlinks (goo.gle, hubs.ly, bit.ly, mzl.la) for campaign attribution and third-party schedulers (Sprinklr, Buffer).

  1. Two Concrete Proposals to Bring to the standard.site Community / Maintainers:
    • Proposal A (Redirect-Following Origin Verification in cardyb / Link Extractors): When a shortlink (https://goo.gle/``<slug>) returns an HTTP 301/302 redirect to a destination page whose <link rel="canonical" href="https://opensource.googleblog.com/..."> matches the resolved destination origin, link extractors (cardyb.bsky.app) should evaluate /.well-known/site.standard.publication and <link rel="site.standard.document"> against the resolved canonical origin rather than aborting verification because the shortlink hostname differs.

    • Proposal B (AppView Backlink Indexing by Canonical URI / site.standard.document): Document in the standard.site implementation guide how AppViews and enterprise publishing tools should normalize <link rel="canonical"> and associatedRefs when posts are published via enterprise social schedulers that do not call cardyb.

Or, am I just missing a way to handle this that I haven’t thought of yet. I’d love to get your feedback.

5 Likes

ohhhhh … this is super interesting! I am going to point some people at this.

3 Likes

Proposal A sounds very reasonable to me.

I work at Bluesky and have passed this along to the app team internally.

2 Likes

I think the problem with Proposal A is that a destination of a shortened url can be changed. Although our builtin shortener doesn’t do this, but our customers might be using bit.ly etc which does allow changing the destination, and hence our link extractor might be creating a record that doesn’t match.
Proposal B, I am interested into what could go in that guide, we can certainly add to the metadata but again, if not at run time might suffer the same problem as A

On a related note, I think trust of shortened url is important, right now it looks ugly no matter what: 1) user gets a warning if url was different 2) shortener link is displayed which would reduce the chance for clicks (Where will I end up)

3 Likes

This is a very important point that must not be overlooked. Not because it is likely to happen a lot, but because of the importance of trust in the system and the possibility of bad actors misusing this in some way.

2 Likes

One can add ./well-known/…. on the shorter domain to declare if the shortener allow changing destination or not, although that can be abused too.

A safer approach would be having a repo of trusted shorteners (those who don’t allow changing the url)