https://issues.apache.org/jira/browse/CASSANDRA-21720
On Fri, 25 Sept 2026 at 21:09, Štefan Miklošovič <[email protected]> wrote: > > I do not have time nor knowledge to do a full-fledged "what is API" CEP. > > Based on this discussion, I will gate the visibility of system_metrics > behind a system property for now, by default it will not be visible. > > I believe this is the minimum we can do with maximum effect of not > exposing this when it is obviously unstable / the subject of further > improvements. It does not seem to me we are confident enough to say > that system_metrics virtual keyspace is stable. > > > On Thu, Sep 10, 2026 at 2:20 AM Yifan Cai <[email protected]> wrote: > > > > FYI, the CEP-65 thread > > (https://lists.apache.org/thread/s8dsl4j1n95vpqtgxxzs20phs2g39m5j) is > > discussing the same API lifecycle question for the shared utils library > > instead of vtables. > > > > Same tension showing up in two places suggests the community wants a clear > > API contract in general, not just for metrics. Might be worth one > > project-wide policy instead of deciding it twice. > > > > - Yifan > > > > On Wed, Sep 9, 2026 at 12:47 PM Maxim Muzafarov <[email protected]> wrote: > >> > >> As an example from the OpenSearch ecosystem, plugins marked with the > >> @Experimental annotation must be explicitly enabled in the YAML > >> configuration and are disabled by default. This also makes their use > >> quite visible from an operations perspective, since the configuration > >> is easy for devops to grep when they need to understand the current > >> cluster state or investigate a problem. > >> > >> ENABLE DEBUGGING / ENABLE METRICS is also a reasonable approach. > >> However, it may be slightly less visible operationally because it is a > >> runtime operation rather than something recorded in the static > >> configuration (even if it's stored in TCM). > >> > >> Example from my recent experience: > >> > >> @Experimental > >> public class IpfixSource implements Source<Record<Event>> > >> > >> experimental: > >> enabled_plugins: > >> source: > >> - IpfixSource > >> > >> On Wed, 9 Sept 2026 at 13:18, Bernardo Botella > >> <[email protected]> wrote: > >> > > >> > I would like to echo what David mentions here of the exposed vtables > >> > being exposed to be considered part of the public APIs, therefore the > >> > expectation should be that anyone can start building on top of them with > >> > some guarantees of them not being broken. > >> > > >> > I also agree that, part of handling those expectations, can be to have > >> > some of those APIs marked as experimental/prone to change/use them at > >> > your own risk. For that, I don’t know if relying only on adding a > >> > EXPERIMENTAL word to the comments field should be the way to go. Maybe > >> > we need to just take a step back and consider them in the same boat we > >> > consider any other API, which should have similar expectations and > >> > guarantees that are consistent across the project. > >> > > >> > Forgetting for a moment that these APIs are vtables, I would think that > >> > for all APIs across the project we need to be able to specify the > >> > contract, and whether they are stable, experimental, or any other > >> > terminology that help those using them making informed decisions. Also, > >> > like with any other APIs, we need a process to be able to deprecate them. > >> > > >> > I have been trying to look for something project wide, and I don’t think > >> > we have an actual policy around APIs. The closest I could find is this > >> > (1) discussion from 2024 around experimental flagging. I think it would > >> > be good to retake this conversation and end up with some good policies > >> > for APIs the project can follow. > >> > > >> > Now, coming back to the EXPERIMENTAL flag for the vtables, if that’s > >> > what we want to do “per policy” for an API, then we can discuss the best > >> > way to do so? > >> > > >> > https://lists.apache.org/thread/9ptpkd3yymy5wok137c5jytysw373v52 > >> > > >> > > >> > From: Štefan Miklošovič <[email protected]> > >> > Date: Wednesday, 9 September 2026 at 07:27 > >> > To: [email protected] <[email protected]> > >> > Subject: Re: [DISCUSS] Future of system_metrics virtual keyspace > >> > > >> > > If they can see it then it should be assumed fair game; so maybe block > >> > > them from seeing it until they agree to a set of shared rules? > >> > > >> > Yes, this is how I see it too. I think we are in a position that we > >> > might declare all current vtables as API-stable - of course minus > >> > Accord ones as Benedict mentioned, and we can make metrics > >> > experimental as well. > >> > > >> > I scanned how vtables were evolving and we never removed anything, it > >> > really is "stable" in that regard, what we have ever done was that we > >> > were only adding new columns, never removing them. Maybe in one case > >> > we changed the type of a column but otherwise it was addition only. > >> > > >> > So, now we declare that Accord + Metrics are experimental and hide > >> > them until they are not. > >> > > >> > How the declaration of this should look like: > >> > > >> > 1) adding a comment into their CQL schemas that this is experimental > >> > and probably subject of change > >> > 2) a user would need to explicitly enable them via system properties, > >> > your ENABLE METRICS is not a bad idea per se but I think that it is > >> > just a "syntactic sugar" and a system property is just enough at this > >> > point. > >> > 3) we document this, in NEWS.txt, that this set of tables are > >> > experimental and hidden. I think that is quite fair and enough, it is > >> > expected that a user is reading this documentation, or at least > >> > should. > >> > > >> > Once they are not in an experimental state anymore, we promote them to > >> > be stable. It would need to be further clarified what that actually > >> > means - what are the deprecation rules around this or if we have to > >> > support it forever. > >> > > >> > On Tue, Sep 8, 2026 at 7:37 PM David Capwell <[email protected]> wrote: > >> > > > >> > > > My view is that virtual tables should not automatically be treated > >> > > > as API-stable. They are for debugging / operator interaction, and > >> > > > intend to expose internal implementation-specific state that is > >> > > > liable to change across minors. > >> > > > >> > > If you expose them via JMX then we fail the build if you break the API > >> > > as we need a stable API for operators… vtables are exposed in a much > >> > > more “public” space so im not sure why the opposite should hold true. > >> > > > >> > > It also doesn’t make sense to me. If we can break the API in a patch > >> > > release and ignore our deprecation process, then operators can’t use > >> > > the APIs… then who are we building them for? In the example that > >> > > started this thread there was desire to stop scraping JMX and use a > >> > > vtable, but if we can break vtables when we feel like it then it would > >> > > be dangerous to depend on the table so we logically should stick to > >> > > JMX as its stable… then why do we have the vtable to begin with (I am > >> > > not saying to drop the table, im just using it as a example against > >> > > the argument)? > >> > > > >> > > We have similar issues with CMS/TCM tables… there was a desire to ask > >> > > clients to move to the TCM peers table as its less queries on startup > >> > > and more up-to-date… but if vtables are unstable we need to rely on > >> > > the old ways. For tablets work we need a new vtable to expose the > >> > > data placement, but again if vtables are unstable then its illogical > >> > > for clients to touch it which becomes a blocker for tablets work... > >> > > > >> > > > We therefore must either avoid exposing internal state and > >> > > > significantly hamper their utility, or else we must reject API > >> > > > compatibility. > >> > > > >> > > This has been brought up before and I do agree with having an ability > >> > > to denote that a table isn’t a stable API, and I would strongly agree > >> > > with a solution to allow this. But this must be clear to a user else > >> > > it’s an impossible situation for everyone. > >> > > > >> > > If you look at other projects you have flags you can call to get > >> > > access to internal state and experimental features, you can create a > >> > > CEP for this as it would be a way to have experimental and unstable > >> > > apis > >> > > > >> > > cqlsh> ENABLE DEBUGGING; > >> > > cqlsh> select * from system_unstable.why_is_repair_acting_up; > >> > > > >> > > You could even use this to allow specific feature > >> > > > >> > > cqlsh> ENABLE METRICS; > >> > > cqlsh> select * from system_unstable.metrics. > >> > > > >> > > If we hide tables by default and expose a way to opt-in, we can define > >> > > rules around them. We could have tables in experimental and be free > >> > > to change them every major until we harden the API, in which case we > >> > > promote it to the top level API. We could also use this to create > >> > > tables we never intent to make stable; stuff that leaks internal state > >> > > so it changes with that internal state. > >> > > > >> > > > My preference would be to standardise on a project policy of > >> > > > defaulting virtual tables to API unstable unless explicitly declared > >> > > > as stable (for programmatic access) > >> > > > >> > > And how would users ever discover this? They use CQL to find tables, > >> > > they see it has the data they need, they build automation using it… > >> > > then we break them when they upgrade and loose their trust. > >> > > > >> > > From a user’s point of view, why should they care about the > >> > > implementation details of a table? If it’s disk backed vs in-memory > >> > > why should they care? If they can see it then it should be assumed > >> > > fair game; so maybe block them from seeing it until they agree to a > >> > > set of shared rules? > >> > > > >> > > > >> > > > On Sep 8, 2026, at 6:41 AM, Štefan Miklošovič > >> > > > <[email protected]> wrote: > >> > > > > >> > > > I agree and also do not think that vtables should be automatically > >> > > > treated as API-stable, should be "case by case" as yours are. > >> > > > > >> > > > But some vtables seem to be queried / parsed already and it will > >> > > > cause > >> > > > a breakage as commentators on CASSANDRA-21539 reported, not sure how > >> > > > to go about it, if we should codify which vtables are considered > >> > > > API-stable and which are experimental, maybe just by putting its > >> > > > experimental status into CQL table description (into "comment") or by > >> > > > gating it behind a system property or similar. > >> > > > > >> > > > If metrics vtables are to be changed and we do not want to cause > >> > > > confusion we might hide them by default and turn it off as Accord has > >> > > > it. > >> > > > > >> > > > On Tue, Sep 8, 2026 at 3:03 PM Benedict Elliott Smith > >> > > > <[email protected]> wrote: > >> > > >> > >> > > >> My view is that virtual tables should not automatically be treated > >> > > >> as API-stable. They are for debugging / operator interaction, and > >> > > >> intend to expose internal implementation-specific state that is > >> > > >> liable to change across minors. We therefore must either avoid > >> > > >> exposing internal state and significantly hamper their utility, or > >> > > >> else we must reject API compatibility. > >> > > >> > >> > > >> To avoid an earlier argument about this, Accord virtual tables are > >> > > >> simply off by default, so that the user must read the commentary > >> > > >> that they are not API stable when enabling them. It may be > >> > > >> suboptimal for users to realise this mid-incident though, but I > >> > > >> cannot promise API compatibility for deep internal state that may > >> > > >> cease to exist entirely. > >> > > >> > >> > > >> My preference would be to standardise on a project policy of > >> > > >> defaulting virtual tables to API unstable unless explicitly > >> > > >> declared as stable (for programmatic access), and - if it makes > >> > > >> some people happy - to report a client warning on first access to > >> > > >> such a table. > >> > > >> > >> > > >> > >> > > >> > >> > > >> On 2026/09/08 12:51:01 Maxim Muzafarov wrote: > >> > > >>> Hi Stefan, > >> > > >>> > >> > > >>> Thank you for bringing this topic up. I think there is still some > >> > > >>> room > >> > > >>> for improvement here. > >> > > >>> > >> > > >>> Out of the options you outlined, I don't think the first option > >> > > >>> excludes the second one. My preference would be to improve the API > >> > > >>> in > >> > > >>> the 6.0 release, while still keeping it experimental for at least > >> > > >>> one > >> > > >>> major release. This would give us some time to collect feedback > >> > > >>> from > >> > > >>> real-world usage before treating the API as stable. > >> > > >>> > >> > > >>> We already had a discussion about experimental virtual tables and > >> > > >>> the > >> > > >>> rules around them: > >> > > >>> > >> > > >>> [DISCUSS] Adding experimental vtables and rules around them > >> > > >>> https://lists.apache.org/thread/xlv5rodt9v77rzrqssp3p63yjg0b88v4 > >> > > >>> > >> > > >>> A few additional thoughts from my side: > >> > > >>> > >> > > >>> 1. > >> > > >>> If we want to change the UX or the schema of the metrics virtual > >> > > >>> tables, I think it is better to do the larger changes in 6.0. This > >> > > >>> way, users could expect only smaller or mostly cosmetic changes in > >> > > >>> later releases rather than a complete redesign. > >> > > >>> > >> > > >>> At the same time, I think the expected usage patterns need to be > >> > > >>> defined more clearly. Saying that there are "usability issues" does > >> > > >>> not give us much direction unless we describe how users are > >> > > >>> expected > >> > > >>> to query, filter, and export metrics. Otherwise, some of the > >> > > >>> proposed > >> > > >>> changes may become a matter of preference rather than solving a > >> > > >>> concrete problem. > >> > > >>> > >> > > >>> 2. > >> > > >>> There is also a pending improvement waiting for a reviewer, which > >> > > >>> did > >> > > >>> not receive the attention it deserved: > >> > > >>> https://issues.apache.org/jira/browse/CASSANDRA-19666 > >> > > >>> > >> > > >>> This change improves the efficiency of bulk metrics exports > >> > > >>> (benchmarks are attached to the issue). The issue also describes in > >> > > >>> more detail the problems with keeping all metrics in a single large > >> > > >>> collection, especially when querying or exporting them in bulk. > >> > > >>> > >> > > >>> I would therefore prefer to address the larger known UX and schema > >> > > >>> issues in 6.0, while still keeping the API in an experimental state > >> > > >>> for some time afterwards. This would give us more flexibility to > >> > > >>> make > >> > > >>> smaller adjustments based on actual usage, rather than committing > >> > > >>> too > >> > > >>> early to the current schema. > >> > > >>> > >> > > >>> On Mon, 7 Sept 2026 at 12:24, Štefan Miklošovič > >> > > >>> <[email protected]> wrote: > >> > > >>>> > >> > > >>>> There is quite a thread on dev Slack discussing the > >> > > >>>> implementation of > >> > > >>>> virtual tables for metrics. Recently, there was also a lot of > >> > > >>>> discussion on (2) where we initially removed a column in a virtual > >> > > >>>> table which received a lot of pushback saying that it is a public > >> > > >>>> API > >> > > >>>> and we can not remove a column like that. > >> > > >>>> > >> > > >>>> Okay, fair enough. > >> > > >>>> > >> > > >>>> But when thinking about what was written in Slack about metrics > >> > > >>>> and > >> > > >>>> how vtables are modeled, I think that it is currently pending to > >> > > >>>> be > >> > > >>>> re-modelled quite a lot. We will very likely need to do this if we > >> > > >>>> want to support e.g. more UX-friendly querying and scrapping of > >> > > >>>> these > >> > > >>>> metrics. Various people (Caleb, David) seem to be not satisfied > >> > > >>>> with > >> > > >>>> how system_metrics look like. (Please correct me if I am wrong). > >> > > >>>> > >> > > >>>> However, when 6.0 is out and we do not do anything about this, > >> > > >>>> then by > >> > > >>>> what was said in 21539 means that we will not be able to > >> > > >>>> restructure > >> > > >>>> vtable metrics and we will need to have yet another set of metric > >> > > >>>> vtables on top of what we ship? > >> > > >>>> > >> > > >>>> What happens with virtual metrics tables after 6.0 when we will > >> > > >>>> need > >> > > >>>> to (massively) re-work their schemas? > >> > > >>>> > >> > > >>>> Possible options: > >> > > >>>> > >> > > >>>> 1) focus on its restructuralization before 6.0 is out because > >> > > >>>> after > >> > > >>>> that it will be effectively set in stone as that is our public API > >> > > >>>> 2) mark it as experimental and object of further re-modelling > >> > > >>>> 3) keep it as it is and never change it > >> > > >>>> 4) ??? > >> > > >>>> > >> > > >>>> Regards > >> > > >>>> > >> > > >>>> (1) https://the-asf.slack.com/archives/CK23JSY2K/p1716324924466269 > >> > > >>>> (2) https://issues.apache.org/jira/browse/CASSANDRA-21539 > >> > > >>> > >> > >
